Method and apparatus for the reproduction of high-order Ambisonics audio signals

By introducing space distortion processing into high-end Ambisens audio technology, the problem of inadequate audio playback at different screen sizes is solved, and the synchronous matching of audio and video is achieved, improving the audience experience.

JP7678917B2Active Publication Date: 2025-05-16DOLBY INTERNATIONAL AB
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024135173
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-03-06
Filing Date
2024-08-14
Publication Date
2025-05-16
Estimated Expiration
2033-03-05

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to video playback of different screen sizes, resulting in the playback of spatial audio that does not adapt to changes in the video screen, affecting the audience's experience.

Method used

The space distortion process is combined with the high-order Ambisens (HOA) audio format to adapt to video playback of different screen sizes. The specific method includes encoding metadata of the reference screen size and performing distortion processing of the audio scene according to the target screen size when decoding.

Benefits of technology

It realizes flexible adaptation of spatial audio content, ensures that the playback position of the audio object matches the position of the video object, and improves the synchronization experience between audio and video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678917000005
    Figure 0007678917000005
  • Figure 0007678917000006
    Figure 0007678917000006
  • Figure 0007678917000007
    Figure 0007678917000007
Patent Text Reader

Abstract

To provide a method and a device, and a non-transitory computer-readable medium for decoding a higher-order ambisonics (HOA) signal that lets a spatial audio content represented as a coefficient of sound filed dissolution match the position where a sound reproduction position of an object on a screen is visible with the corresponding eyes by adapting the spatial audio content to video screens of various size in a combination with video reproduction.SOLUTION: In a device, an HOA-represented signal from a storage device 92 stored with an OA=encoded signal is HOA-decoded by an HOA decoder 93, enters a renderer 95 through distortion means 94, and is output as a speaker signal 91 for a set of speakers. The distortion means receives reproduced adaptive information 90, and provides it so as to adapt the decoded HOA signal suitably. The reproduced adaptive information is derived from differences in width, height and curvature between the original screen and present screen.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method and device for the reproduction of an original high-order Ambisonics audio signal assigned to a video signal to be presented on a current screen but which was originally generated for a different screen. [Background technology]

[0002] One way to store and process the three-dimensional sound field of a spherical microphone array is the Higher-Order Ambisonics (HOA) representation. Ambisonics uses orthonormal spherical functions to describe the sound field at and in a region around a reference point in space, also known as the origin or sweet spot. The accuracy of such a description is determined by the Ambisonics order N, where a finite number of Ambisonics coefficients describe the sound field. The maximum Ambisonics order for a spherical array is limited by the number of microphone capsules, which is the number of Ambisonics coefficients O=(N+1). 2 The advantage of such an Ambisonics representation is that the sound field reproduction can be individually adapted to almost any given speaker position arrangement. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] EP1518443B1 [Patent Document 2] EP1318502B1 [Patent Document 3] PCT / EP2011 / 068782 [Patent Document 4] EP11305845.7 [Patent Document 5] EP11192988.0 [Non-patent literature]

[0004] [Non-Patent Document 1] Sandra Brix, Thomas Sporer, Jan Plogsties, "CARROUSO - An European Approach to 3D-Audio", Proc. of 110th AES Convention, Paper 5314, 12-15 May 2001, Amsterdam, The Netherlands [Non-Patent Document 2] Ulrich Horbach, Etienne Corteel, Renato S. Pellegrini and Edo Hulsebos, "Real-Time Rendering of Dynamic Scenes Using Wave Field Synthesis", Proc. of IEEE Intl. Conf. on Multimedia and Expo (ICME), pp.517-520, August 2002, Lausanne, Switzerland. [Non-Patent Document 3] Richard Schultz-Amling, Fabian Kuech, Oliver Thiergart, Markus Kallinger, "Acoustical Zooming Based on a Parametric Sound Field Representation", 128th AES Convention, Paper 8120, 22-25 May2010, London, UK [Non-Patent Document 4] Franz Zotter, Hannes Pomberger, Markus Noisternig, "Ambisonic Decoding With and Without Mode-Matching: A Case Study Using the Hemisphere", Proc. of the 2nd International Symposium on Ambisonics and Spherical Acoustics, 6-7 May 2010, Paris, France Summary of the Invention [Problem to be solved by the invention]

[0005] While facilitating a flexible and universal representation of spatial audio almost independent of the speaker setup, its combination with video playback on screens of various sizes can become cumbersome as the spatial sound reproduction is not adapted accordingly.

[0006] Stereo and surround sound are based on discrete speaker channels and there are very specific rules about where to place the speakers in relation to the video display. For example, in a theater environment, the center speaker is located in the center of the screen and the left and right speakers are located to the left and right of the screen. Thereby, the speaker setup inherently scales with the screen, i.e. for small screens the speakers are closer to each other and for large screens the speakers are further apart. This has the advantage that the sound mixing can be done in a very coherent way, i.e. sound objects related to the visible objects on the screen can be reliably localized between the left, center and right channels. Thus the listener's experience matches the creative intent of the sound artist from the mixing stage.

[0007] But these advantages are also the disadvantages of channel-based systems: the flexibility to change speaker placement is very limited. This disadvantage increases as the number of speaker channels increases. For example, 7.1 and 22.2 formats require precise placement of each individual speaker, making it extremely difficult to adapt audio content for non-optimal speaker positions.

[0008] Another drawback of channel-based formats is that the ability to pan sound objects between the left, center, and right channels is limited due to precedence effects, especially for large listening setups such as in a theater environment. For off-center listening positions, panned audio objects may "walk" into the speaker closest to the listener. Many films have therefore been mixed with important screen-related sounds, especially dialogue, mapped only to the center channel, which results in a very stable localization of those sounds on the screen, but at the cost of a less-than-optimal overall sound scene spread.

[0009] A similar compromise is typically chosen for the rear surround channels: since the precise locations of the speakers reproducing those channels are largely unknown in the production, and the density of those channels is rather low, typically only ambient and uncorrelated items are mixed into the surround channels, thereby reducing the chance of noticeable reproduction errors in the surround channels, but at the cost of not being able to faithfully place discrete sound objects anywhere other than on the screen (or even in the center channel, as discussed above).

[0010] As mentioned above, the combination of spatial audio with video playback on screens of various sizes can be annoying as the spatial sound playback is not adapted accordingly. Depending on whether the actual screen size matches the screen size used in the production or not, the direction of the sound object may deviate from the direction of the visible object on the screen. For example, if the mix is ​​performed in an environment with a small screen, the sound object associated with the screen object (e.g. the voice of an actor) will be localized within a relatively narrow cone seen from the mixer position. If this content is mastered to a sound field based representation and played in a theater environment with a much larger screen, there will be a noticeable mismatch between the wide field of view to the screen and the narrow cone of the sound object relative to the screen. A large mismatch between the position of the visible image of the object and the position of the corresponding sound is annoying to the viewer and therefore seriously affects the perception of the film.

[0011] More recently, parametric or object-oriented representations of audio scenes have been proposed, which describe an audio scene by a composition of individual audio objects with a set of parameters and properties. For example, in non-patent literature 1, 2, object-oriented scene descriptions have been proposed, primarily for wave-field synthesis systems.

[0012] US 6,313,633 describes two different approaches to address the problem of adapting audio playback to the visible screen size. The first approach determines the playback position for each sound object individually, depending on parameters such as its direction and distance to a reference point as well as aperture angles and the positions of both the camera and projection equipment. In practice, a tight coupling between the visibility of such objects and the associated sound mixing is not typical. On the contrary, some deviation of the sound mix from the relevant visible object may in fact be tolerated for artistic reasons. Furthermore, it is important to make a distinction between direct and ambient sound. Last but not least, the incorporation of physical camera and projection parameters is rather complicated and such parameters are not always available. The second approach (see claim 16) describes a pre-calculation of sound objects based on the above procedure, but assumes a screen of a fixed reference size. This scheme requires a linear scaling of all position parameters (in Cartesian coordinates) to adapt the scene to a screen larger or smaller than the reference screen. However, this means that the virtual distance to the sound objects also doubles when accommodating a screen of twice the size. This is a mere "breathing" of the acoustic scene without any change in the angular position of the sound objects relative to the listener at the reference seat (i.e. the sweet spot). With this approach it is not possible to generate listening results that are faithful to changes in the relative size of the screen in angular coordinates (aperture angle).

[0013] Another example of an object-oriented sound scene description format is described in US Pat. No. 5,399,663. Here, the audio scene contains, besides the various sound objects and their properties, information about the properties of the room to be reproduced as well as the horizontal and vertical opening angles of the reference screen. At the decoder, similar to the principle of US Pat. No. 5,399,663, the position and size of the actual available screen are determined and the reproduction of the sound objects is individually optimized to match the reference screen.

[0014] For example, in US Pat. No. 5,399,544, sound-field oriented audio formats such as Higher Order Ambisonics HOA have been proposed for a universal spatial representation of sound scenes, and in terms of recording and playback, sound-field oriented processing offers a good trade-off between universality and practicality, since, like object-oriented formats, it can be scaled to virtually any spatial resolution. On the other hand, several straightforward recording and generation techniques exist that allow to derive a natural recording of the real sound field, as opposed to a fully synthetic representation required for object-oriented formats. Obviously, sound-field oriented audio content does not contain information about individual sound objects, so the mechanisms introduced above for adapting object-oriented formats to different screen sizes are not applicable.

[0015] To date, only a few publications are available that describe means of manipulating the relative positions of individual sound objects contained in a sound-field-oriented audio scene. The family of algorithms described in [3] requires the decomposition of the sound field into a limited number of discrete sound objects. The position parameters of these sound objects can then be manipulated. This approach has the drawback that the audio scene decomposition is error-prone and any errors in the discrimination of the audio objects are likely to lead to artifacts in the sound rendering.

[0016] Many publications, such as the abovementioned non-patent literature 1 and non-patent literature 4, are concerned with optimizing the playback of HOA content for "flexible playback layouts". These techniques address the problem of using irregularly spaced speakers, but none of them aim to change the spatial composition of the audio scene. [Means for solving the problem]

[0017] The problem solved by the present invention is to adapt spatial audio content, represented as coefficients of a sound field decomposition, to video screens of different sizes, such that the sound reproduction positions of objects on the screen match their corresponding visible positions. This problem is solved by a method as disclosed in claim 1. An apparatus utilizing this method is disclosed in claim 2.

[0018] The present invention allows the reproduction of spatial sound field oriented audio to be systematically adapted to the reproduction of its linked visible objects, thereby fulfilling a significant prerequisite for faithful reproduction of spatial audio for cinema.

[0019] According to the present invention, sound field oriented audio scenes are adapted to different video screen sizes by applying a space warping process as disclosed in US Pat. No. 6,233,639 in combination with sound field oriented audio formats as disclosed in US Pat. No. 6,233,639 and US Pat. No. 6,233,639. One advantageous process is to encode the reference size of the screen (or the viewing angle from a reference listening position) used in the content creation and transmit it as metadata together with the content.

[0020] Alternatively, a fixed reference screen size is assumed during encoding and decoding, and the decoder knows the actual size of the target screen. The decoder distorts the sound field in such a way that all sound objects in the direction of the screen are compressed or stretched according to the ratio of the size of the target screen to the size of the reference screen. This can be achieved, for example, using a simple two-segment piecewise linear warping function as described below. In contrast to the above state of the art, this stretching is essentially limited to the angular position of the sound items and does not necessarily lead to a change in the distance of the sound objects to the listening area.

[0021] Below are described several embodiments of the invention that allow control over which parts of the audio scene are manipulated or not.

[0022] In principle, the method of the invention is suitable for the reproduction of an original high-order Ambisonics audio signal that has been assigned to a video signal that is to be presented on a current screen but was originally generated for a different screen. decoding the higher order Ambisonics audio signal to provide a decoded audio signal; receiving or establishing playback adaptation information derived from differences in width and possibly height and possibly curvature between said original screen and said current screen; adapting the decoded audio signal by warping it in the spatial domain, the playback adaptation information controlling the distortion such that for a current screen viewer and for a listener of the adapted decoded audio signal, a perceived position of at least one audio object represented by the adapted decoded audio signal matches a perceived position on the screen of an associated video object; Rendering and outputting the adapted decoded audio signal to a speaker.

[0023] In principle, the device according to the invention is suitable for reproducing an original high-order Ambisonics audio signal that has been assigned to a video signal that is to be presented on a current screen but was originally generated for a different screen. means adapted to decode the Higher Order Ambisonics audio signal to provide a decoded audio signal; · means adapted to receive or establish reconstruction adaptation information derived from differences in width and possibly height and possibly curvature between said original screen and said current screen; means adapted to adapt the decoded audio signal by warping it in the spatial domain, the playback adaptation information controlling the distortion such that for a current screen viewer and for a listener of the adapted decoded audio signal, a perceived position of at least one audio object represented by the adapted decoded audio signal matches a perceived position on the screen of an associated video object; and means for rendering and outputting the adapted decoded audio signal to a speaker.

[0024] Advantageous further embodiments of the invention are disclosed in the respective dependent claims. [Brief description of the drawings]

[0025] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1] FIG. 1 illustrates an exemplary studio environment. [Diagram 2] FIG. 1 illustrates an exemplary movie theater environment. [Diagram 3] FIG. 2 is a diagram illustrating a distortion function f(φ). [Figure 4] FIG. 13 is a diagram illustrating a weighting function g(φ). [Diagram 5] FIG. 13 shows original weights. [Figure 6] FIG. 13 illustrates weights after warping. [Figure 7] FIG. 2 illustrates a warping matrix. [Figure 8] FIG. 1 illustrates a known HOA process. [Figure 9] FIG. 2 illustrates a process according to the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0026] Figure 1 shows an exemplary studio environment with a reference point and a screen, and Figure 2 shows an exemplary cinema environment with a reference point and a screen. Different projection environments lead to different opening angles of the screen when viewed from the reference point. With state-of-the-art sound field oriented reproduction techniques, audio content produced in a studio environment (opening angle 60°) does not match the screen content in a cinema environment (opening angle 90°). To allow the content to adapt to the different characteristics of the reproduction environment, it is necessary to transmit the opening angle of 60° in the studio environment together with the audio content.

[0027] For clarity, these drawings simplify the situation to a 2D scenario.

[0028] In high-order Ambisonics theory, the spatial audio scene is expressed as the coefficients A of the Fourier-Bessel series: n m For a source-free volume, the sound pressure is described as a function of spherical coordinates (radius r, tilt angle θ, azimuth angle φ) and spatial frequency k=ω / c, where c is the speed of sound in air.

[0029]

number

[0030] The spatial composition of an audio scene can be distorted by a technique disclosed in US Pat. No. 5,399,663.

[0031] The relative positions of sound objects contained within a two-dimensional or three-dimensional Higher Order Ambisonics HOA representation of an audio scene can be varied, where the magnitude O in Input vector A with in determines the coefficients of the Fourier series of the input signal, and the magnitude O out Output vector A with out Determine the Fourier series coefficients of the output signal that have been changed accordingly. in is the inverse matrix Ψ1 of the mode matrix Ψ1 -1 Using s in =Ψ1 -1 A in For regularly-spaced speaker positions, we compute the input signal s in The input signal s in A out =Ψ2s in The output vector A of the adapted output HOA coefficients is warped and encoded in the spatial domain by computing out Here, the mode vector of the mode matrix Ψ2 is the angle of the original speaker position to the output vector A out The angle is then modified according to a distortion function f(φ) that maps one-to-one to the target angle of the target speaker position in

[0032] The speaker density correction is done by applying a gain weighting function g(φ) to the virtual speaker output signal s in Applying it to the result s out In principle, any weighting function g(φ) can be specified. One particularly advantageous variant is one that is proportional to the derivative of the distortion function f(φ).

number

[0033]

number

[0034] Decoding, weighting and warping / decoding are usually done with a size of O warp ×O warp Transformation matrix T=diag(w)Ψ2diag(g)Ψ1 -1 Here, diag(w) denotes a diagonal matrix with the values ​​of the window vector w as its diagonal elements, and diag(g) denotes a diagonal matrix with the values ​​of the gain function g as its diagonal elements. out ×O in To obtain the transformation matrix T, we use the spatial warping operation A out =TA in To perform the transformation, the corresponding columns and / or rows of the transformation matrix T are removed.

[0035] Figures 3 to 7 show spatial distortion in the two-dimensional (circular) case, showing an example piecewise linear distortion function for the scenario in Figures 1 / 2 and its effect on the panning function of 13 example regularly arranged loudspeakers. The system stretches the front sound field by a factor of 1.5 to accommodate the larger screen in the cinema. Sound items coming from other directions are therefore compressed. The distortion function f(φ) resembles the phase response of a discrete-time all-pass filter with a single real-valued parameter and is shown in Figure 3. The corresponding weighting function g(φ) is shown in Figure 4.

[0036] FIG. 7 illustrates a 13×65 single-step transform warping matrix T. The log-magnitude values ​​of the individual coefficients of this matrix are shown by grayscale or shading type based on the accompanying grayscale or shading bar. This example matrix has N orig = 6 input HOA orders and N warp = 32. This higher output order is needed to capture most of the information that is spread by the transformation from lower order coefficients to higher order coefficients.

[0037] A useful property of this particular warping matrix is ​​that a significant portion of it is zero, which allows to save a great deal of computational power when implementing this operation.

[0038] Figures 5 and 6 show the distortion characteristics of beam patterns generated by several plane waves. Both figures result from the same 13 input plane waves with the same amplitude "1" at φ positions 0, 2 / 13π, 4 / 13π, 6 / 13π, ..., 22 / 13π and 24 / 13π, resulting in 13 different angular amplitude distributions, i.e., the overdetermined regular decoding operation s=Ψ -1The resulting vector s of A is shown. Here, the HOA vector A is either the original or a distorted version of the set of plane waves above. The numbers outside the circle represent the angle φ. The number of virtual speakers is much larger than the number of HOA parameters. The amplitude distribution or beam pattern for a plane wave coming from the front direction is located at φ=0.

[0039] Figure 5 shows the weight and amplitude distributions of the original HOA representation. All 13 distributions have a similar shape and are characterized by the same width of the main lobe. Figure 6 shows the weight and amplitude distributions for the same sound object, but after the distortion operation has been performed. The object is now farther away from the front direction at φ=0 degrees, and the main lobes around the front direction are wider. These modifications of the beam pattern result in a higher order N distribution of the distorted HOA vector. warp = 32. A mixed-order signal is generated with the local order varying over space.

[0040] The distortion characteristic (f(φ in In order to derive the HOA coefficients, additional information is sent or provided. For example, the following characteristics of the reference screens used in the mixing process can be included in the bitstream: -Direction to the center of the screen, ·width, · Reference screen height. These are all expressed in polar coordinates measured from the reference listening position (also called the "sweet spot").

[0041] Additionally, the following parameters may be required for special applications: The shape of the screen, e.g. flat or spherical, Screen distance, Information about maximum and minimum visible depth in case of stereoscopic 3D video projection.

[0042] Those skilled in the art will know how such metadata can be encoded.

[0043] In the following, we assume that the encoded audio bitstream contains at least the three parameters mentioned above, namely the orientation of the center of the reference screen, the width and the height. For simplicity, we further assume that the center of the actual screen is identical to the center of the reference screen, e.g. directly in front of the listener. We further assume that the sound field is represented only in a 2D format (not in a 3D format), and that tilt changes for this are ignored (e.g. when the selected HOA format does not represent a vertical component, or when the sound editor decides that the mismatch between the picture and the tilt of the on-screen sound source is small enough to be unnoticeable to the casual observer). The transition to arbitrary screen positions and the 3D case is straightforward for those skilled in the art. We further assume, for simplicity, that the screen configuration is spherical.

[0044] With these assumptions, only the screen width can vary between content and actual setup. In the following, a suitable two-segment piecewise linear distortion characteristic is defined. The actual screen width is given by the aperture angle 2φ w,a (i.e., φ w,a represents one-sided angle). The reference screen width is the angle φ w,r and is part of the meta-information delivered in the bitstream. For faithful reproduction of a sound object in the forward direction, i.e. on the video screen, every position (in polar coordinates) of the sound object is defined by a factor φ w,a / φ w,r Conversely, all sound objects in the other direction are moved according to the remaining space. The distortion characteristic leads to the following results:

[0045]

number

[0046] Furthermore, if the chosen HOA representation also provides for tilt, and the sound editor considers the vertical angle subtended by the screen to be of interest, then the angular height of the screen, θ h (one-sided height) and related factors (e.g., the ratio of actual height to reference height θ h,a / θ h,r ) can be applied to the slope as part of the warping operator.

[0047] As another example, assuming a flat screen in front of the listener rather than a spherical one may require more elaborate distortion characteristics than in the examples above. Again, this can concern width-only or width+height distortion.

[0048] The above exemplary embodiment has the advantage of being fixed and fairly simple to implement. On the other hand, it does not allow any control of the adaptation process from the production side. The following embodiment introduces the process for more control in a different way. EXAMPLES

[0049] Example 1: Separation between screen-related sounds and other sounds There can be various reasons why such control techniques are needed. For example, it can be advantageous to manipulate direct sounds differently from ambient sounds, rather than all of the sound objects in the audio scene being directly connected to objects visible on the screen. This distinction can be performed by scene analysis on the rendering side. However, it can be significantly improved and controlled by adding additional information to the transmitted bitstream. Ideally, the decision of which sound items should be adapted to the actual screen characteristics - and which ones should be left untouched - should be left to the artist doing the sound mixing.

[0050] Different ways to convey this information to the rendering process are possible.

[0051] ●Two full sets of HOA coefficients (signals) are defined in the bitstream, one describing objects related to the visible item, the other representing independent or surrounding sounds. At the decoder, only the first HOA signal undergoes adaptation to the actual screen geometry, the other one is left untouched. Before playback, the manipulated first HOA signal and the unmodified second HOA signal are combined.

[0052] As an example, a sound engineer may decide to mix screen-related sounds such as dialogue or individual sound effects items into one signal and ambient sounds into a second signal, so that the ambient sounds always remain the same no matter which screen is used for the playback of the audio / video signal.

[0053] This type of processing has the additional advantage that the HOA orders of the two constituent sub-signals can be individually optimized, whereby the HOA order for the screen-related sound objects (i.e. the first sub-signal) is higher than the one used for the ambient signal components (i.e. the second sub-signal).

[0054] ● Via flags attached to the time-space-frequency tiles, the mapping of the sound is defined as being screen-related or independent. For this purpose, the spatial properties of the HOA signals are determined, for example via a plane wave decomposition. Each of the spatial domain signals is then input to a time segmentation (windowing) and a time-frequency transformation. Thereby, a three-dimensional set of tiles is defined, which can be individually marked, for example by a binary flag stating whether the content of that tile should be adapted to the actual screen geometry or not. This sub-embodiment is more efficient than the previous one, but it limits the flexibility of defining which parts of the sound scene should or should not be manipulated. EXAMPLES

[0055] Example 2: Dynamic Adaptation In some applications it may be necessary to change the signaled reference screen characteristics in a dynamic way. For example, the audio content may be the result of concatenating repurposed content segments from different mixes. In this case, the parameters describing the reference screen parameters change over time and the adaptation algorithm is dynamically modified, i.e. for every change of the screen parameters the applied distortion function is recalculated accordingly.

[0056] Another application example arises from mixing different HOA streams prepared for different sub-parts of the final visible video and audio scene, where it is advantageous to allow more than two (or more than two in the case of example 1 above) HOA signals, each with their individual screen characteristics, in a common bitstream. EXAMPLES

[0057] Example 3: Alternative Implementation Instead of distorting the HOA representation prior to decoding via a fixed HOA decoder, information about how to adapt the signal to the actual screen can also be integrated into the decoder design. This implementation is an alternative to the basic realization described in the exemplary embodiment above. However, it does not change the signaling of screen characteristics in the bitstream.

[0058] In Figure 8, the HOA encoded signal is stored in storage device 82. For presentation in a cinema, the HOA rendered signal from device 82 is HOA decoded in HOA decoder 83, passed through renderer 85, and output as speaker signals 81 for a set of speakers.

[0059] In Figure 9, the HOA encoded signal is stored in storage device 92. For presentation, for example in a cinema, the HOA rendered signal from device 92 is HOA decoded in HOA decoder 93, passes through distortion stage 94 to renderer 95 and is output as speaker signals 91 for a set of speakers. Distortion stage 94 receives playback adaptation information 90 as described above and uses this to adapt the decoded HOA signal accordingly.

[0060] A few additional notes are included. [Appendix 1] 1. A method for reproducing an original high-order Ambisonics audio signal assigned to a video signal to be presented on a current screen but originally generated for a different screen, the method comprising: decoding the higher order Ambisonics audio signal to provide a decoded audio signal; receiving or establishing playback adaptation information derived from differences in width and possibly height and possibly curvature between said original screen and said current screen; adapting the decoded audio signal by warping it in the spatial domain, the playback adaptation information controlling the distortion such that for a current screen viewer and for a listener of the adapted decoded audio signal, a perceived position of at least one audio object represented by the adapted decoded audio signal matches a perceived position on the screen of an associated video object; rendering and outputting the adapted decoded audio signal for a speaker; method. [Appendix 2] 2. The method of claim 1, wherein the higher-order Ambisonics audio signal includes a plurality of audio objects assigned to corresponding video objects, and to a viewer and listener of the current screen, an angle or distance of the audio objects is different from the respective angle or distance of the video objects on the original screen. [Appendix 3] 3. The method of claim 1 or 2, wherein a bitstream carrying the original high-order Ambisonics audio signal also includes the playback adaptation information. [Appendix 4] 4. The method of any one of claims 1 to 3, wherein in addition to the warping, weighting by a gain function is performed to obtain a resulting uniform sound amplitude per opening angle. [Appendix 5] 5. The method of any one of claims 1 to 4, wherein two full coefficient sets of high-order Ambisonics audio signals are decoded, a first audio signal representing an object related to a visible object and a second audio signal representing an independent or peripheral sound, wherein only the first decoded audio signal is subjected to adaptation to the actual screen geometry by distortion and the second decoded audio signal is left untouched, and wherein the adapted first decoded audio signal and the non-adapted second decoded audio signal are combined before playback. [Appendix 6] 6. The method of claim 5, wherein the HOA orders of the first and second audio signals are different. [Appendix 7] 7. The method of any one of claims 1 to 6, wherein the playback adaptation information is dynamically changed. [Appendix 8] 1. A reproduction device for an original high-order Ambisonics audio signal assigned to a video signal to be presented on a current screen but originally generated for a different screen, said device comprising: means adapted to decode the Higher Order Ambisonics audio signal to provide a decoded audio signal; · means adapted to receive or establish reconstruction adaptation information derived from differences in width and possibly height and possibly curvature between said original screen and said current screen; means adapted to adapt the decoded audio signal by warping it in the spatial domain, the playback adaptation information controlling the distortion such that for a current screen viewer and for a listener of the adapted decoded audio signal, a perceived position of at least one audio object represented by the adapted decoded audio signal matches a perceived position on the screen of an associated video object; means for rendering and outputting the adapted decoded audio signal to a speaker; Device. [Appendix 9] 9. The apparatus of claim 8, wherein the higher-order Ambisonics audio signal includes a plurality of audio objects assigned to corresponding video objects, and to a viewer and listener of the current screen, an angle or distance of the audio objects is different from the respective angle or distance of the video objects on the original screen. [Appendix 10] 10. The apparatus of claim 8 or 9, wherein a bitstream carrying the original high-order Ambisonics audio signal also includes the playback adaptation information. [Appendix 11] 11. The apparatus of any one of claims 8 to 10, wherein in addition to the distortion, weighting by a gain function is performed to obtain a resulting uniform sound amplitude per opening angle. [Appendix 12] 12. The apparatus of any one of claims 8 to 11, wherein two full coefficient sets of high-order Ambisonics audio signals are decoded, a first audio signal representing objects related to a visible object and a second audio signal representing independent or peripheral sounds, wherein only the first decoded audio signal is subjected to adaptation to the actual screen geometry by distortion and the second decoded audio signal is left untouched, and wherein the adapted first decoded audio signal and the non-adapted second decoded audio signal are combined before playback. [Appendix 13] 13. The apparatus of claim 12, wherein the HOA orders of the first and second audio signals are different. [Appendix 14] 14. The apparatus of any one of claims 8 to 13, wherein the playback adaptation information is dynamically changed. [Appendix 15] 1. A method of generating digital audio signal data, the method comprising: providing data of an original high-order Ambisonics audio signal to be assigned to the video signal; providing reproduction adaptation information data derived from the width and possibly the height and possibly the curvature of a screen on which said video signal can be presented, the playback adaptation information data can be used to adapt a decoded version of the Higher Order Ambisonics audio signal by distortion in the spatial domain such that for a viewer of the video signal and a listener of the adapted decoded audio signal on a current screen having a width different from that of the original screen, the perceived position of at least one audio object represented by the adapted decoded audio signal matches the perceived position of an associated video object on the current screen, method.

Claims

1. 1. A method for decoding an encoded Higher Order Ambisonics (HOA) signal describing a sound field, comprising: decoding the encoded HOAs to obtain a first set of decoded HOA signals representing dominant components of the sound field and a second set of decoded HOA signals representing ambient components of the sound field; combining the first set of decoded HOA signals and the second set of decoded HOA signals to produce a combined set of decoded HOA signals; determining a transformation matrix for warping the combined set of decoded HOA signals, the transformation matrix being based on a production screen size and a target screen size, the transformation matrix being further based on a diagonal matrix of speaker correction gains; method.

2. A non-transitory computer readable medium containing instructions that, when executed by a processor, perform the method of claim 1.

3. 1. An apparatus for decoding an encoded Higher Order Ambisonics (HOA) signal describing a sound field, comprising: an audio decoder for decoding the encoded HOA to obtain a first set of decoded HOA signals representing dominant components of the sound field and a second set of decoded HOA signals representing ambient components of the sound field; a combiner that combines the first set of decoded HOA signals and the second set of decoded HOA signals to produce a combined set of decoded HOA signals; a processor for determining a transformation matrix for warping the combined set of decoded HOA signals, the transformation matrix being based on a production screen size and a target screen size, the transformation matrix being further based on a diagonal matrix of speaker correction gains. Device.

Citation Information

Patent Citations

  • Method for coding audio

    EP1318502B1

  • Device and method for determining a reproduction position

    EP1518443B1

  • Metapneumovirus strains and their use in vaccine formulations and as vectors for expression of antigenic sequences and methods for propagating virus

    EP2494986A1

  • Method and apparatus for changing the relative positions of sound objects contained within a higher-order ambisonics representation

    EP2541547A1

  • Playing spatially shaped audio

    JP2002505058A