Distributed Conversational Binaural Rendering

JP2025517658A5Pending Publication Date: 2026-05-15DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2023-05-09
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing solutions for real-time audio rendering based on listener orientation and position require high data transmission bandwidth and processing power, leading to increased power consumption and latency, which is challenging to mitigate without compromising the quality of the audio experience.

Method used

A method involving a first processing module that generates multiple audio presentations associated with different listener orientations and positions, and a second processing module that receives conversion parameters and user orientation data to modify the main presentation, reducing the need for high-bandwidth data transmission and minimizing latency.

Benefits of technology

This approach allows for efficient, low-latency rendering of object-based audio content that appears to be fixed in space, even with limited processing capabilities, thereby enhancing the user's immersive audio experience while reducing power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, a system, and a computer program product for processing audio. The method includes receiving at least one input audio signal and generating a main rendered presentation and an additional rendered presentation, each rendered presentation being associated with a listener orientation and / or position. The method further includes determining conversion parameters for converting the main rendered presentation into the additional rendered presentation and determining a deviation value based on the user's orientation and / or position and the listener orientation and / or position. The method further includes determining modified conversion parameters based on the conversion parameters and the deviation value and applying the modified conversion parameters to the main rendered presentation to generate an output presentation associated with the user's orientation and / or position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications This application claims the priority of U.S. Provisional Patent Application No. 63 / 340,181, filed on May 10, 2022, the entire content of which is incorporated herein by reference.

[0002] Technical field of the invention The present invention relates to a method for distributed rendering of audio signals.

Background Art

[0003] Binaural audio content is becoming increasingly popular in the form of stereo audio signals, for example, for reproduction on headphones or speaker systems with crosstalk cancellation. For example, object - based audio content can be rendered as a binaural stereo presentation for headphones using a Head - related Transfer Function (HRTF). Object - based audio content includes one or more audio objects associated with arbitrarily time - varying positions in three - dimensional space. For example, an audio object may be intended to be perceived by a listener as an audio object to the right of the listener, above the listener, or moving along an orbit around the listener. Thus, object - based audio can provide an acoustic effect that enhances the listener's sense of immersion.

[0004] HRTFs have been developed that describe the interaural time difference, the interaural level difference, reflections that occur in the human ear, and the frequency response of the human ear as a function of the orientation and / or position of the listener's head. Using such HRTFs, binaural audio signals can be generated for any static or dynamic placement of audio objects in three-dimensional space. Additionally, room reflections and / or reverberation are typically added to create a sense of perceived distance and space.

[0005] In some cases, the rendering of object-based audio content is adapted substantially in real time based on the orientation and / or position of the listener such that instead of fixing the audio objects to the listener's head, they are fixed to the environment. Thus, when the listener moves their head, the sound image is shifted accordingly, and the rendering is adapted to make the listener perceive that the audio objects are fixed in space rather than to their head. As an example, a listener is presented with an audio presentation that is rendered such that an audio object is perceived to be located to the right of the listener. If the listener changes orientation and faces the opposite direction, this change in orientation is recorded by an orientation detector, which provides this information to a renderer, which then modifies the rendering to provide a modified presentation that is rendered such that the audio object is perceived to be located to the left of the listener. The effect of this is that the audio object seems as if it is fixed in the listener's environment, allowing the listener to move and / or change orientation within this space. This form of orientation and / or position-modified rendering is sometimes referred to as interactive binaural rendering and is particularly useful in game applications, extended reality (XR) applications, augmented reality (AR) applications, and virtual reality (VR) applications. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0006] The drawback of existing solutions for audio rendering based substantially on the orientation and / or position of a listener in real time is that the rendering is associated with high requirements for data transmission bandwidth and processing power, which increases the power consumption of the device performing the rendering. At the same time, in order to enable the rendering of a convincing audio image that includes audio objects that are not fixed to the listener's head, but rather appear to be fixed in space or moving along a fixed trajectory in space, it is important to keep the latency, i.e., the time delay between the listener changing the orientation and / or position of the head and the associated modification in the audio presentation, very low, typically on the order of tens of milliseconds.

[0007] Accordingly, a first problem is to provide a rendering process based on orientation and / or position that provides a sufficiently low latency and responds quickly to any change in the orientation and / or position of the listener. Since a latency on the order of 17 ms may be perceptible to many listeners, the latency between a change in orientation and / or position and the presentation of a modified audio presentation to the listener should ideally be substantially less than 100 ms. However, such low latency is practically difficult to achieve due to the inherent delay introduced by the rendering process itself, as well as the (typically wireless) transmission of sensor and audio data from the orientation tracking device worn by the user and the system, service or computer configured to perform the audio rendering.

[0008] To reduce latency, the orientation and / or position tracking device, audio renderer, and loudspeaker may be integrated into the same wearable device (e.g., earphones or a VR headset). However, this still presents a second problem related to the computational power required for substantially real-time orientation / position-based rendering that responds quickly to changes in the listener's orientation and / or position, and the associated high power consumption. Object-based audio may include a number of assets representing the ambient situation, point sound sources, sound effects, dialog, and other important elements, all of which need to be rendered in real time in response to changes in the listener's orientation and / or position that can occur suddenly and be very rapid (e.g., due to the listener quickly changing orientation, looking up or down, or walking within the environment). Wearable devices such as VR headsets, smart glasses, earphones, or glasses generally do not have the processing power or battery capacity required to sustain this audio rendering for very long. Thus, in many applications, the orientation and / or position information is transmitted from the wearable device to a more powerful companion device such as a phone, tablet, computer, game console, or cloud computer (e.g., an edge server) that performs the rendering, and the rendered presentation is returned to the wearable device. However, the communication between the companion device and the wearable device significantly increases latency, especially when the communication occurs over a common wireless communication channel such as Bluetooth® that can introduce significant latency.

[0009] To achieve a very low latency, more capable wearable devices can be used with improved processing performance and, for example, a larger battery. However, a third challenge arises to physically accommodate the improved device capabilities. This is because wearable devices become larger and inconvenient to use (e.g., they become bulkier and / or heavier to accommodate the necessary processing, power, and cooling components). In general, the bandwidth for communicating with wearable devices is also limited, and since multiple audio elements in object-based audio content require a significant amount of bandwidth, some audio elements may need to be removed or compressed, which degrades the Quality of Experience (QoE). Since it is difficult to achieve sufficient bandwidth in wireless communication, some solutions rely on a wired data connection to the wearable device, which greatly hinders the flexibility of the wearable device and makes it difficult to use outdoors or for the user to move around freely.

[0010] An object of the present disclosure is to present a method for rendering audio content, particularly object-based audio content, that responds substantially in real time to changes in the listener's orientation and / or position and overcomes or at least mitigates the problems of the conventional solutions identified above.

Means for Solving the Problems

[0011] According to a first aspect of the present invention, in a first processing module, receiving at least one input audio signal and, in the first processing module, generating a main rendered presentation and an additional rendered presentation, each rendered presentation being associated with a first and a second listener orientation and / or position respectively, there is provided a method of processing audio. The method further includes, in the first processing module, determining conversion parameters for converting the main rendered presentation into the additional rendered presentation, and, in a second processing module, receiving the conversion parameters and the main rendered presentation generated by the first processing module. The method further includes, in the second processing module, receiving user orientation and / or position data indicating the user's orientation and / or position, and, in the second processing module, determining an orientation and / or position deviation value based on the user's orientation and / or position and the first and second listener orientations and / or positions, and, in the second processing module, determining a modified conversion parameter based on the conversion parameter and the orientation and / or position deviation value, and, in the second processing module, applying the modified conversion parameter to the main rendered presentation to generate an output presentation associated with the user's orientation and / or position.

[0012] That is, the first processing module pre-emptively renders at least two presentations associated with different listener orientations and / or positions and, for each presentation except one (the main presentation), determines conversion parameters that can be used to convert the main presentation into the at least one additional rendered presentation.

[0013] The "orientation" of a listener or user means the rotational orientation of the head of the intended listener or user. For example, the orientation can be defined by one or more of the pitch angle, yaw angle, and roll angle. The "position" of a listener or user means the position of the head of the listener or the head of the user in one or more of the front / back, left / right, and up / down directions. For example, the position may be defined by a Cartesian coordinate system having perpendicular X, Y, and Z axes. It is understood that different listener orientations and / or positions may differ in one of the orientation and position, or may differ in both the orientation and position. Some implementations assume that only changes in orientation (having 1, 2, or 3 degrees of freedom) are considered, while in other implementations only changes in position (having 1, 2, or 3 degrees of freedom) are considered.

[0014] The orientation and / or position deviation value may be a linear or non-linear distance between two orientations and / or positions. Further, the orientation and / or position deviation may be a perceptually weighted distance between two orientations and / or positions, as will be explained in more detail below.

[0015] The conversion parameters may be updated for each time-frequency tile of the time-frequency representation. As will be explained below, for an audio representation having two channels, each set of conversion parameters may include four or five conversion parameters (some of which may be complex-valued), or even just two real-valued conversion parameters, which constitutes an amount of data that can be transmitted quickly with low latency. The conversion parameters are still sufficient to accurately describe the orientation / position conversion from the main presentation to an additional presentation, and can be used to find modified conversion parameters (e.g., using interpolation) if the user's orientation / position does not correspond to the orientation / position associated with the additional presentation.

[0016] Thus, even if the conversion parameters are updated frequently, for example, for each time-frequency tile, the conversion parameters can be efficiently transmitted to the second processing module, representing only a small amount of data (compared to hundreds or thousands of samples for representing time-frequency tiles of an audio channel).

[0017] Furthermore, the application and / or modification of the conversion parameters is computationally efficient and can be executed quickly even on a processing module with limited processing capabilities, which means that the second processing module can be implemented on limited devices such as headphones, earphones, wireless earphones, true wireless earphones, smart glasses or VR / AR / XR headsets. By receiving the rendered main presentation associated with the first listener orientation / position and the conversion parameters associated with the second listener orientation / position, the second processing module can quickly modify the conversion parameters and apply them to the main presentation to shift the presentation to the second listener orientation / position when the second listener orientation / position better matches the actual user orientation / position. Also, in order to more accurately follow the user's orientation / position, it is also possible to modify the conversion parameters, for example, using interpolation, before applying the conversion parameters to the main presentation.

[0018] Using this method, the rendering of the input audio signal can be shifted based on the user's orientation / position, whereby the user is presented with an audio presentation that appears to be fixed in space. As an illustrative example, an audio asset is associated with music coming from a virtual stage directly in front of the listener, and the user listens to these audio assets using earphones while standing in a physical space. When the user turns their head to the right, the rendering is adjusted so that the listener is presented with an audio presentation that makes it seem as if the music is coming from the left. This is an example of modifying the presentation to follow the user's orientation with respect to the virtual three-dimensional space of the audio asset. If the listener moves towards or away from the virtual stage, the user can be presented with an audio presentation where the music becomes louder or softer. This is an example of modifying the presentation to follow the user's position with respect to the virtual three-dimensional space of the audio asset. One or more audio assets may also include audio objects that move along a trajectory within the virtual three-dimensional space. By shifting the rendering of the audio asset based on the user's orientation / position, it is possible to provide the listener with an audio presentation that makes it seem as if the trajectory along which the audio object moves is fixed in the virtual three-dimensional space.

[0019] In some implementations, the first and second listener orientations and / or positions are different yaw orientations at respective first and second pitch orientations, and the method further includes, in a second processing module, obtaining reduced transformation parameters associated with a third pitch orientation, the reduced transformation parameters being configured to transform a main rendered presentation or an additional rendered presentation into a pitch-changed rendered presentation having the third pitch orientation, and, in the second processing module, applying the reduced transformation parameters to the main rendered presentation based on an orientation deviation to generate an output presentation.

[0020] That is, each set of conversion parameters may be associated with a respective orientation where the yaw (where the user looks left or right) is different at a given pitch angle (where the user looks up or down), and the conversion parameters capture a very prominent inter-aural effect for changing yaw angles. On the other hand, in order to span different pitch angles, a set of reduced conversion parameters having fewer parameter values (e.g., one actual gain value per channel) compared to the (non-reduced) conversion parameters is transmitted for each yaw angle for a plurality of pitch angles deviating from a given pitch angle. Thus, by taking into account that the sensitivity to an audio presentation shift in yaw is different from the presentation shift in pitch, it is possible to reduce the amount of information transmitted to the second processing module without degrading the quality of experience (QoE).

[0021] According to a second aspect of the present invention, there is provided a computer program product including instructions which, when executed by a computer, cause the computer to execute the method according to the first aspect.

[0022] According to a third aspect of the present invention, there is provided a system comprising a first processing module communicating with a second processing module, the first and second processing modules being configured to execute the method according to the first aspect.

[0023] The computer program product and system according to the second and third aspects feature the same or equivalent advantages as the method according to the first aspect. Any function described with respect to the method may have a corresponding feature in the system or computer program product, and vice versa.

Brief Description of the Drawings

[0024] Aspects of the present invention will be described in more detail with reference to the accompanying drawings showing presently preferred embodiments.

[0025]

Figure 1

[0026]

Figure 2

[0027]

Figure 3

[0028]

Figure 4

[0029]

Figure 5

[0030]

Figure 6

[0031]

Figure 7

[0032]

Figure 8

[0033]

Figure 9

[0034]

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0035] The systems and methods disclosed in the present application may be implemented as software, firmware, hardware, or combinations thereof. In a hardware implementation, the division of tasks does not necessarily correspond to the division into physical units. Conversely, one physical component may have multiple functions, or one task may be executed by several physical components working together.

[0036] Computer hardware may be, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, AR / VR wearable, an automotive infotainment system, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that computer hardware. Further, the present disclosure relates to any set of computer hardware that executes instructions, individually or jointly, to perform any one or more of the concepts described herein.

[0037] Some or all components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code containing a set of instructions that, when executed by one or more processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify the actions to be taken is included. Thus, one example is a typical processing system (e.g., computer hardware) that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem that includes a hard drive, an SSD, RAM, and / or ROM. A bus subsystem may be included for communication between components. Software may be present in the memory subsystem and / or within the processor during its execution by the computer system.

[0038] One or more processors may operate as a stand-alone device or may be connected to other processors, e.g., network-connected. Such a network may be built on various different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0039] Software can be distributed on computer-readable media that can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes various forms of physical (non-transitory) storage media such as, but not limited to, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Further, communication media (transitory) typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media, as is well known to those skilled in the art.

[0040] FIG. 1 shows a distributed rendering system 1 according to some implementations. The distributed rendering system 1 has three subsystems 11, 13, 15. More specifically, the distributed rendering system 1 includes a multi-presentation renderer module 11, a multi-presentation encoder 13, and an interactive renderer 15. At least one of the subsystems 11, 13, 15 is implemented in a device separate from the device implementing at least one of the other subsystems 11, 13, 15, which means that the complete rendering process is distributed across at least two devices that communicate with each other. Two of the subsystems 11, 13, 15 may be implemented on the same device, and the remaining subsystems 11, 13, 15 are implemented by separate devices.

[0041] As described below, the amount of input data and the computational complexity of the processes executed vary between different subsystems. The advantage of the distributed rendering system 1 of FIG. 1 is that the amount of data transmitted to the interactive renderer 15 is minimized, while the interactive renderer 15 is associated with the least complex processes among the three subsystems. This makes the interactive renderer 15 suitable for implementation in computationally limited and power-constrained devices such as wearable devices, while the other two subsystems 11, 13 can be implemented in more computationally capable devices such as smartphones, computers, or game consoles that communicate with the device implementing the interactive renderer 15.

[0042] Thus, in some implementations, the multi-presentation renderer 11 and the multi-presentation encoder 13 are implemented on a high-performance device (or optionally, on two different high-performance devices that communicate with each other), the interactive renderer 15 is implemented on a separate constrained device, and the high-performance device is configured to communicate with the constrained device. Examples of high-performance devices may include smartphones, tablets, computers (e.g., desktops or laptops), game consoles, cloud computers, or servers. Examples of constrained devices may include a pair of headphones, earphones, wireless earphones, smart glasses, true wireless earphones, or VR / AR / XR headsets. It may be beneficial for the constrained device to communicate with the high-performance device using a wireless connection (e.g., WiFi or Bluetooth®), although it is also envisioned that the communication may occur via a wired connection.

[0043] The processes executed by the multi-presentation renderer 11, the multi-presentation encoder 13, and the interactive renderer 15 are described in further detail with reference to FIG. 1.

[0044] The multi-presentation renderer 11 is configured to render at least two audio presentations based on one or more audio assets 10. The audio presentations are labeled as R 1 , …, R p , …, R P , which means that the multi-presentation renderer 11 generally renders P presentations, where P ≥ 2. Each of at least two presentations R 1 , …, R P is associated with a different listener orientation and / or position with respect to the audio asset 10. The term “listener orientation and / or position” is used to indicate the assumed listener orientation / position with respect to the audio asset 11.

[0045] The audio asset 10 may often include one or more spatialized audio objects, often simply called audio objects. An audio object is an audio signal associated with spatial attributes such as a position or an incident direction in a three-dimensional space. How one or more audio objects should be rendered to form an audio presentation depends on the assumed listener orientation and / or position with respect to the audio object.

[0046] The multi-presentation renderer 11 selects a plurality of possible listener orientations / positions labeled as V 1 , V 2 , …, V P for the audio asset 10, and for each of the plurality of listener orientations / positions V 1 , V 2 , …, V P , it renders the individual presentations R 1 , …, R P . For example, for the plurality of listener orientations / positions V 1 , V 2 , …, V Pis selected to span a range of orientations (indicated by pitch, yaw, and roll angles) and / or positions (indicated by Cartesian coordinates X, Y, Z) in the three-dimensional space of the audio asset. Note that the multi-presentation renderer 11 can select the listener's orientation / position regardless of the user's actual measured orientation / position. That is, the multi-presentation renderer 11 is V 1 ,V 2 ,…,V P rendering a plurality of possible presentations corresponding to listeners oriented in, but generally none of these positions exactly correspond to the actual user orientation V L .

[0047] For example, each audio presentation R 1 ,…,R P is a pair of binaural audio signals extracted using respective HRTFs, and the orientation / position of the HRTF for the audio asset 10 varies between the respective HRTFs.

[0048] That is, the multi-presentation renderer 11 obtains at least two orientations V 1 ,…,V p ,…,V P and for each orientation, renders the corresponding presentation R 1 ,…,R p ,…,R P based on the audio asset 10. In some implementations, as schematically shown in FIG. 2, the orientations / positions V 1 ,…,V p ,…,V P span different combinations of pitch and yaw angles at a given point in the three-dimensional space of the audio asset 10. For example, the orientation V 1 indicates 0 degrees of pitch and 0 degrees of yaw, the orientation V 2 indicates 0 degrees of pitch and 5 degrees of yaw, the orientation V 3 indicates 5 degrees of yaw and -5 degrees of yaw, etc. Similarly, the orientation / position can be selected to span a variety of X, Y, Z positions.

[0049] each orientation / position V 1 ,…,V P for each, a separate audio presentation R 1 ,…,R P is rendered. To achieve this, the multi-presentation renderer 11 may comprise a plurality of renderers 12a, 12b, 12c configured to render presentations R 1 ,…,V P each associated with an individual orientation / position V 1 ,…,V P and presentations R 1 ,…,R P associated based on the orientation / position V 1 ,…,R P and the audio asset 10. In some implementations, each presentation R

[0050] is a binaural audio presentation suitable for playback on headphones including two audio channels, a left audio channel and a right audio channel. However, it is envisioned that the presentation may be of other types such as a monaural presentation, a stereo presentation, or a surround presentation (e.g., 5.1 or 7.1 presentation).

[0050] The multi-presentation renderer 11 renders at least two presentations R 1 and R 2 . Generally, the multi-presentation renderer 11 is beneficial when rendering a large number of presentations such as at least 10 presentations (P≧10), at least 20 presentations (P≧20), or at least 50 presentations (P≧50) spanning a wide area of the listener orientation / position, and / or ensuring that the distance between two listener orientations / positions is not too large.

[0051] The multi-presentation renderer 11 transmits the rendered presentations R 1 ,…,R P to the multi-presentation encoder 13.

[0052] The multi-presentation encoder 13 receives P presentations R 1 ,…,RP Receive all of them, and for all presentations except one, determine a set of conversion parameters W p . That is, the multi-presentation encoder 13 designates one of the P presentations as the main presentation, and for all the remaining (at least one) presentations, determines the associated conversion parameters. The remaining presentations are called additional presentations. Hereinafter, without loss of generality, presentation R 1 is the main presentation, that is, presentation R 2 , …, R P is the additional presentation R 2 , …, R P , and the associated conversion parameters W 2 , …, W P are determined for each of the remaining presentations R 2 , …, R P .

[0053] Each set of conversion parameters W p is configured to convert the main presentation R 1 in the listener orientation V 1 to the presentation R p at the position V p , where the index p ranges from 2 to P and P ≥ 2.

[0054] To determine the conversion parameters W p , the multi-presentation encoder 13 includes one or more parameter generators 14b, 14c, and each parameter generator 14b, 14c receives as input two presentations, namely, the main presentation R 1 and one of each of the additional presentations R 2 , …, R P . Each parameter generator 14b, 14c generates conversion parameters W 1 for converting the main presentation R P to each of the respective additional presentations R P .

[0055] The operation of the parameter generator 14b and the properties of the conversion parameters will be described in detail below. It is understood that other types of conversion parameters can be determined and used in the same manner, and that other parameter generators 14a can operate in exactly the same manner as parameter generator 14b.

[0056] As described above, the parameter generator 14b receives two rendered presentations, namely R 1 and the labeled main presentation, and an additional rendered presentation R p . The two presentations R 1 , R p are in the same format. For example, the main presentation R 1 and the additional rendered presentation R p can both be binaural audio signals, both be stereo audio signals, both be mono audio signals, or both be surround audio signals (e.g., 5.1 signals). In the following exemplary implementation, it is assumed that the formats of the presentations R 1 , R p are binaural formats including two audio channels, but it should be noted that the same process can be similarly executed for presentations in other formats.

[0057] Each rendered presentation R 1 , R p includes a left channel and a right channel (forming a pair of binaural signals). Thus, for the main presentation R 1 and the additional presentation R p , the following equation holds.

Equation

[0058] To determine the transformation matrix ^M P the parameter generator may determine an initial least-squares solution ^M p by minimizing the root mean square error between the presentation R P and the main presentation R 1 to which ^M P is applied. That is, the initial least-squares solution ^M P can be determined as [Number] and can be determined as such.

[0059] The initial least-squares solution ^M P from Equation 4 can be determined iteratively. Alternatively, the initial least-squares solution ^M P can be expressed as a closed-form solution. For example, the covariance matrix of the channels of the presentation, for example [Number] is represented as R y,p,p and the correlation matrix between two different presentations R p1 and R p2 is represented as R y,p1,p2 and the closed-form solution of Equation 6 is

Number

Number

[0060] In some implementations, a more accurate reconstruction of the additional presentation R p is desired, whereby an additional diagonal gain matrix G and / or decorrelation contribution gain g d,p is determined and the transformation matrix ^MP is modified to obtain an improved modified transformation matrix M p which is used. For example, ^M P even if it is the solution of Equation 4, its application to the main presentation R 1 may, in some cases, result in a covariance matrix different from the covariance matrix of the presentation R p The covariance matrix of the predicted presentation ^R p is given by R ^y,p,p and may be different from the covariance matrix R [Number] and can be expressed as. Here, R ^y,p,p is the covariance matrix R p of the actual presentation R y,p,p which may be different. This deviation in the covariance matrix can be corrected by the decorrelation coefficient scaled by the diagonal gain matrix G and / or the decorrelation gain g d,p Specifically, an improved reconstruction model of R from the main presentation R 1 is established as follows p * [Number] [Number] Here, the function Ψ(·) represents a decorrelator that generates an input signal from an output signal according to the following two requirements [Number] Here, 〈·〉 represents the expected value operator

[0061] From Equation 8, the following is shown. Assuming g d,p ≦0, the resulting covariance matrix R p * of this improved prediction R z,p,p is given by the following equation [Number]

[0062] Thus, the diagonal gain matrix G and gd,p By setting appropriate values of this improved prediction R p * covariance matrix R z,p,p is the true presentation R p covariance matrix R y,p,p can be guaranteed to be equal to. R z,p,p = R y,p,p The solution to this problem, i.e., = R, can be shown to be found as follows.

Number

[0063] Thus, the improved transformation matrix M p is

Number

[0064] In some implementations, the procedure for determining the transformation matrix ^M p or the improved transformation matrix M p and the decorrelation gain g d,p is repeated for each time - frequency tile of the presentations R 1 and R p . The parameter generator 14b determines the initial least - squares solution ^M p or the improved transformation matrix M p and the decorrelation gain g d,p for each time - frequency tile of each presentation. Thus, the initial least - squares solution ^M p or the improved transformation matrix M p and the decorrelation gain g d,penables an exact reconstruction of the presentation R from R, even if the presentation varies over time and / or frequency. 1 of the presentation R from p

[0065] The elements of the transform matrix ^M p or the improved transform matrix M p should be noted that they may be complex-valued. Furthermore, each element of the transform matrix ^M p or the improved transform matrix M p may be a vector value (for example, forming a 3D matrix), and each element defines a plurality of discrete real or complex filter samples that define a FIR filter.

[0066] In the above example, the presentations R 1 , …, R p are binaural audio signals having two channels. In such an example, ^M p and M p are 2×2 matrices (or 2×2×M matrices, where M is the number of discrete filter bank samples). In general, for a presentation having N channels, ^M p and M p are N×N matrices or N×N×M matrices.

[0067] The transform matrix ^M p or the improved transform matrix M p and the decorrelation gain g d,p are combined into a set of transform parameters W for each presentation and time-frequency tile. For example, for each presentation and time-frequency tile p

Number

[0068] The set of transform parameters W for each presentation p is sent to the interactive renderer 15 together with the main presentation R 1 . The set of transform parameters W p is for the main presentation R​​1 It is understood that much less data is required for transmission as compared. For example, the main presentation R 1 is related to a time-frequency tile representation (e.g., STFT representation) having hundreds or thousands of complex-valued audio samples within a frame. On the other hand, the presentation R p associated with the conversion parameter W p is related to, per presentation per time-frequency tile, W p ={^M p} in the case of four (potentially complex) values, or W p ={M p ,g d,p} in the case of five values (four of which are potentially complex). If all STFT frequency audio samples within each frame are grouped into 5 to 50 tiles or frequency bands for which the conversion parameter W p is calculated, the conversion parameter will include approximately 25 to 250 parameter values, which is one to two orders of magnitude smaller than the audio samples within the main presentation. Apart from the reduction in the number of parameters, the accuracy of the conversion parameter can typically be made significantly lower than the required accuracy of the audio samples, which also contributes to the reduction of information when the transmitted data is digitally quantized. Thus, it is much more efficient to transmit an additional set of the conversion parameter W p compared to the transmission of the additional presentation R p .

[0069] The main presentation R 1 is received by the interactive renderer 15 along with at least one set of the conversion parameter W p associated with the orientation / position V p . Since what the associated orientation / position V p is may not be directly derivable from the conversion parameter W p itself, the orientation / position V p may be explicitly indicated as the vector V p associated with each set of the conversion parameter W p . Alternatively, the orientation / position V1 ,…,V P are pre-determined and may be locally stored in the interactive renderer 15.

[0070] Referring further to FIG. 2, the positions V 2 ,…,V P are schematically shown as to how they are distributed so as to span a range of translational positions and / or rotational orientations. The positions V illustrated in FIG. 2 1 , V 2 , V 3 , V 4 span different translational positions in the XY plane and / or different rotational orientations in the pitch-yaw plane. It is understood that additional orientations / positions may be added in a third dimension (Z-axis or roll axis) perpendicular to the XY plane or pitch-yaw plane shown in FIG. 2. The positions V 1 are associated with the main presentation R 1 , and for the remaining positions V 2 , V 3 , V 4 , associated conversion parameters W 1 are available in the interactive renderer 15 to convert the main presentation R to a presentation associated with one of the positions V 2 , V 3 , V 4 . 2 , W 3 , W 4

[0071] The interactive renderer 15 is the user's orientation / position V LIt also receives user-oriented and / or position data indicating. The listener-oriented / position data can be received from an orientation / position detector such as a head-tracking detector. Examples of orientation / position detectors include magnetic sensors, gyro sensors, GPS receivers, accelerometers, and UV / IR / visible light sensors (e.g., camera sensors). The orientation / position detector may be included in the same device as the interactive renderer 15 to enable low-latency communication between the orientation detector and the interactive renderer 15. For example, the orientation detector and the interactive renderer 15 may be included in a set of earphones or earbuds. Additionally or alternatively, the orientation / position tracker is provided outside the device implementing the interactive renderer 15, and the orientation / position tracker transmits the user-oriented and / or position data to the interactive renderer 15. An exemplary external orientation tracking device is one or more cameras provided in the user's environment that track the orientation / position of the user's head, for example, using motion tracking or face recognition. Another example of an external orientation tracker is that the user wears one or more light-emitting devices that emit IR, UV, or visible light and are tracked by one or more IR, UV, or visible light sensors provided in the environment, and the orientation / position of the light-emitting device is associated with the orientation / position of the user's head.

[0072] As shown in Figure 3, the user orientation / position V L has conversion parameters W 2 for it, W 3 for it, W 4 where the orientation / position V 1 is generated, V 2 for it, V 3 for it, V 4 and may deviate from one or more of them. Generally, the user orientation / position V L is the orientation / position V 1 for it, V 2 for it, V 3 for it, V 4 and may change continuously or with a much finer granularity compared to the granularity of. For example, the user's yaw orientation V 1 of the main presentation R 1and an orientation V for which conversion parameters are available 2 associated presentation R 2 moved his head so as to face a yaw orientation different from the yaw orientation of 2 . For this purpose, the interactive renderer 15 determines the deviation value between the user's orientation / position V L and the main presentation R 1 associated orientation / position V 1 and the conversion parameter set W 2 W 3 W 4 associated orientation / position V 2 V 3 V 4 and at least one of V L and comprises a parameter processor 17 configured to determine a modified conversion parameter W

[0073] Determining the modified conversion parameter W L may include selecting the conversion parameters associated with the listener orientation / position V L that is at a minimum distance (e.g., the closest) to 1 V 2 V 3 V 4 and using these parameters to transform the main presentation R 1 Alternatively, determining the modified conversion parameter W L may include interpolating between at least two sets of the conversion parameters W L associated with the listener orientation / position V 1 V 2 V 3 V 4 in the vicinity of 1 W 2 W 3 W 4

[0074] Two listener orientations / positions V 1 V 2 V 3 V 4 ​The deviation value between, and / or the listener orientation / position V 1 、V 2 、V 3 、V 4 and the user orientation V L The deviation value between can be determined as a linear or non - linear distance between the orientations / positions. For example, each orientation / position V 1 、V 2 、V 3 、V 4 、V L may be represented as a point or vector in a coordinate system, and the deviation value between two orientations / positions V 1 、V 2 、V 3 、V 4 、V L can simply be determined as the Euclidean distance, cosine similarity, haversine distance, etc. between two points or vectors.

[0075] Furthermore, the deviation value between two listener orientations / positions V 1 、V 2 、V 3 、V 4 and / or the deviation value between the listener orientation / position V 1 、V 2 、V 3 、V 4 and the user orientation V L can be a perceptually weighted distance. That is, determining the deviation value between two orientations / positions V 1 、V 2 、V 3 、V 4 、V L can involve weighting different components of the distance (e.g., represented by pitch, yaw, roll angles and / or X, Y, Z distances) differently based on the expected perceptual impact on the rendered presentation.

[0076] For example, as will be explained in detail below, a change in orientation in yaw (i.e., the user looks left or right) may be perceptually more important compared to a change in orientation in pitch (the user looks up or down). Thus, the deviation value between two orientations / positions V 1 、V2 , V 3 , V 4 , V L When determining the deviation value between, any distance in yaw may be weighted more compared to the distance in pitch. For example, the orientation V L of the listener closest to the user's orientation V 1 , V 2 , V 3 , V 4 selection of, has a listener orientation with a similar yaw orientation rather than a similar pitch orientation V 1 , V 2 , V 3 , V 4 of, is prioritized.

[0077] The same also applies, for example, when determining the deviation value for different positions that can be represented in a Cartesian coordinate system having X, Y, and Z directions. Here, a change in position in one direction has a greater perceptual impact on the rendered presentation compared to a change in position in another direction. Thus, distances along different axes can be weighted differently such that the distance along one axis affects the deviation value more significantly compared to the distance along another axis.

[0078] Furthermore, the perceptual weighting may emphasize the distance in orientation rather than the distance in position when determining the deviation value, or vice versa. For example, in many implementations, a change in orientation may have a greater perceptual effect compared to a change in position. Therefore, it may be beneficial to assign a greater weight to the distance in orientation compared to the distance in position. As an example, when an audio asset includes an audio object located at a large distance from the user, a small change in position may have a very small perceptual effect, but a small change in orientation may still have a large perceptual effect. Thus, when determining the deviation value between the user orientation / position and the listener orientation / position V 1 , V 2 , V 3 , V 4 , V

[0079] As described above, the modified conversion parameter W L to be determined is the user-oriented V L and the conversion parameter W 2 、W 3 、W 4 associated with the listener-oriented / position V 2 、V 3 、V 4 and the orientation / position V 1 associated with the main presentation R 1 including determining a deviation value between at least one of them, the deviation value being a measure of linear distance, non-linear distance and / or perceptually weighted distance between the user-oriented V L and the listener-oriented V 1 、V 2 、V 3 、V 4 、V

[0080] Thus, the scales of the axes in FIGS. 2 to 6 may be linear, non-linear, and / or perceptually distorted. For example, the listener-oriented / position V 1 、V 2 、V 3 、V 4 in FIG. 2 may have a much larger linear distance of the pitch angle between V 1 and V 3 than the linear distance of the yaw angle between V 1 and V 2 so that these distances may appear to be approximately the same in a distorted or non-linear axis but may not be uniformly distributed in a linear space.

[0081] Referring to FIG. 4, the conversion parameters associated with the orientation / position V L that is the minimum distance (e.g., the closest) to the listener-oriented / position V 1 、V 2 、V 3 、V 4 will be described here. The listener-oriented / position V L is the other position V 1 、V 2 、V 3Compared to the orientation / position V 2 (i.e., associated with smaller deviation values), the transformation parameter W 2 But, W L =W 2 The closest, best matching orientation / position V 1 , V 2 , V 3 , V 4 Determining and using transformation parameters associated with V is a process that can be done very quickly and efficiently, but there can be noticeable acoustic disturbances when the user changes orientation and different transformation parameters are selected. On the other hand, many transformation parameters may vary depending on the orientation / position V. 1 , V 2 , V 3 , V 4 If the acoustic dispersion is generated for a finer grain size distribution, the prominence of these acoustic disturbances may be mitigated.

[0082] Alternatively, the modified transformation parameters W L Determining the orientation / position V 1 , V 2 , V 3 , V 4 User orientation / position relative to V L Based on different orientations / positions V 1 , V 2 , V 3 , V 4 The transformation parameters W associated with 2 , W 3 , W 4 For example, as shown in FIG. 4, the user orientation / position V L is the orientation / position V 4 and V 2 and so the modified transformation parameters W L Determining V L and V 2 The deviation between L and V 4 Based on the deviation between V and V, 4 The parameter W associated with 4and the parameter W 2 associated with the orientation / position V 2 including interpolating therebetween.

[0083] Thus, the discrete orientation / position V 1 , V 2 , V 3 , V 4 The conversion parameters associated with the orientation / position between are accessible via interpolation such as linear interpolation or using a triangulation interpolation function. The conversion parameter W 1 , W 2 , W 3 , W 4 for each position may all be in the same format. As described above, each set of the conversion parameter W p may indicate the initial least squares solution transformation matrix ^M p , or each set of W p may indicate the improved transformation matrix M p and the decorrelation gain G d,p that can be used to transform the main presentation.

[0084] The main presentation R 1 is associated with the orientation / position V 1 . However, since it may not be necessary to determine any conversion parameters associated with this orientation / position, the parameter processor 17 may associate the position V 1 with the default conversion parameter W 1 , for example W 1 = {M p , g d,p}, M p = 1 and g d,p = 1, allowing the orientation / position V 1 , and the associated default conversion parameter W 1 to be used for interpolation. For example, as shown in FIG. 3, the user orientation / position V L may be between V 1 and V 2 , whereby the interpolation to find W L is between V L and V1 The deviation value between and V L and V 1 Based on the deviation value between and, the orientation / position V 1 is associated with the default parameter W 1 and the orientation / position V 2 is associated with the parameter W 2 is executed by the parameter processor 17 between them.

[0085] In FIGS. 2 and 3, the user orientation / position V L is illustrated as being between two orientations / positions along the DOF axis (e.g., along the X axis or the yaw axis as in FIG. 3), but the user's orientation / position V L can be any orientation / position in a 6DOF system. That is, interpolation can be performed between more than two points. For example, as shown in FIG. 5, the user orientation / position V L is associated with different individual transformation parameters W 1 , W 2 , W 3 , W 4 is associated with a plurality of orientations / positions V 1 , V 2 , V 3 , V 4 is between them, but may be separated from them, and the interpolation is based on the deviation value between the user orientation / position V L and the orientation / position V 1 , V 2 , V 3 , V 4 and is executed between the transformation parameters W 1 , W 2 , W 3 , W 4 .

[0086] Main presentation R 1 and the transformation parameters W 2 , W 3 , W 4It may be transmitted to a second interactive renderer (not shown) implemented on a device separate from the device implementing the first interactive renderer 15. The second interactive renderer performs corresponding processing with the first interactive renderer 15, but is based on the user orientation / position V L2 of the second user. That is, the second interactive renderer determines a second set of modified conversion parameters that can be used to convert the same main presentation R 2 , W 3 , W 4 based on the same conversion parameters W 1 into a presentation associated with the second user orientation / position V L2 . In other words, the distributed rendering system 1 can be used for multicast, and the same audio asset 10 is rendered to a plurality of users via individual interactive renderers 15 that obtain the orientation of each user but use the same multi-presentation renderer 11 and multi-presentation encoder 13.

[0087] (by interpolation or selection of the best matching parameters) After determining the modified conversion parameter W L , the parameter processor 17 provides the conversion parameter W L to the presentation converter 15, and the presentation converter applies the modified conversion parameter W L to the main presentation R 1 to generate an output representation. The output presentation may be referred to as an interactive output presentation. Optionally, the output presentation is provided to one or more loudspeakers (e.g., the loudspeakers of a set of earphones or headphones) that present the output presentation to the user.

[0088] The conversion parameters W 1 , W 2 , W 3 , W 4 are periodically updated for each (optionally partially overlapping) frame and frequency band and provided to the interactive renderer 15. Similarly, the interactive renderer 15 is based on the user orientation / position V LReceives periodic updates and, for each time-frequency tile, determines and applies modified transformation parameter W L to it.

[0089] In some implementations, it may be desirable to use different formats of transformation parameter W p for different DOFs. For example, a listener may not be as sensitive to inaccuracies in a presented rendering when changing orientation / position in some specific DOFs compared to other DOFs. In particular, for rotational orientations, it has been found that listeners are less sensitive to inaccuracies when changing the pitch orientation (e.g., looking up or down) compared to when changing the yaw orientation (e.g., looking from left to right). Regarding changes in pitch, most of the binaural localization cues such as the interaural time difference and level difference that users use to localize sound have been found to remain constant because the perceived position moves along the so-called cone of confusion.

[0090] The most prominent cue for changes in pitch is in the form of changes in the frequency spectrum, while the most prominent cue for changes in yaw is in the form of changing interaural level differences and time delays. Thus, it is assumed that the transformation parameter W p associated with different pitch orientations for a particular yaw orientation is described using fewer parameters compared to a particular yaw orientation.

[0091] For example, the transformation parameter W p as described above may be determined for a set of primary yaw orientations, where the primary yaw orientations have different yaw angles at a given pitch angle (e.g., 0 degrees pitch corresponding to a listener looking horizontally). For at least one of the primary yaw orientations, a set of reduced transformation parameters W pitch is determined for transforming to the pitch orientation (e.g., 15 degrees pitch corresponding to a listener looking up from the horizontal line) associated with the primary yaw orientation at a given pitch. The reduced transformation parameter W pitchincludes, for each frame and frequency band, the gains for each channel in the presentation formation. For example, for a binaural or stereo presentation format, the downsampling transformation parameter W pitch includes, for each time-frequency tile and frequency band, the gain g of the left channel pitch,u,l and the gain g of the right channel pitch,u,r . If a total of U different associated pitch orientations are used for each main yaw direction, the gains of the downsampling transformation parameter are

Number

[0092] Figure 6 schematically shows a plurality of main yaw orientations V p,u shown using format V 1,0 , V 2,0 , where p is the yaw angle index and u is the pitch angle index. Each main yaw direction V 1,0 , V 2,0 is associated with a predetermined pitch (e.g., a pitch of 0 degrees) and each set of transformation parameters W 1 for converting the main presentation R 1 , W 2 . Each set of transformation parameters W 1 , W 2 represents the matrix elements of the initial least-squares solution transformation matrix ^M p for each time-frequency tile, or W pEach set is the main presentation R 1 to the main yaw orientation V 1,0 、V 2,0 and can be used to convert to a presentation associated with the yaw of, an improved conversion matrix M p and decorrelation gain G d,p can be shown.

[0093] For each main yaw orientation V 1,0 、V 2,0 at least one set of reduction transformation parameters W pitch is generated, and the reduction transformation parameters are the main yaw orientation V 1,0 、V 2,0 of the transformation parameters W 1 、W 2 indicates how the transformation parameters W 1 、W 2 should be modified to obtain transformation parameters that convert to presentations with different pitches. For example, the reduction transformation parameters W pitch,2,1 and W pitch,2,-1 are respectively associated with the accompanying pitch orientations V 2,1 and V 2,-1 which are associated with the main yaw orientation V 2 associated with the transformation parameter W 2,0 is associated.

[0094] By comparing Figure 6 with Figure 5, it can be seen that for many pitches and yaw orientations, the reduction transformation parameter W p is replaced by the reduction transformation parameter W pitch is used. Each set of reduction transformation parameters W pitch has only two real-valued gain values, namely g pitch,u,l and g pitch,u,r (or generally one for each channel in the presentation format), so compared to at least four real or complex values for the set of transformation parameters W p , much less data is required to span the same number of orientations, or the same amount of data can represent more orientations. At the same time, the reduction transformation parameter W pitchSince it is used only for pitch conversion where the user is not very sensitive, data reduction can be achieved essentially without degrading the experience quality for the user.

[0095] In one exemplary implementation, the first and second listener-oriented Vs 1,0 V 2,0 are the main listener orientations at respective predetermined first and second pitch orientations (e.g., 0-degree pitch). The first and second listener-oriented Vs 1,0 V 2,0 are associated with respective sets of (complete) conversion parameters W 1 W 2 Optionally, one of the first and second listener-oriented Vs 1,0 V 2,0 is associated with the main presentation and default conversion parameters.

[0096] At least one of the first and second listener orientations is associated with a third listener orientation V 1,0 V 2,0 that has a different pitch orientation from the first or second listener positions V 2,1 but has the same yaw orientation as the first or second listener positions. The third listener orientation V 2,1 is associated with a reduced conversion parameter W 1,0 W 2,0 that has one value per time-frequency tile and channel in the presentation format and can be used to adjust the pitch of the first or second listener orientations V pitch,2,1 The interactive renderer obtains the user-oriented V L and determines an orientation deviation value based on the user-oriented V L and at least one of the first, second, and third listener orientations V 1,0 V 2,0 V 2,1 and applies the reduced conversion parameter W pitch,2,1 based on the orientation deviation value. For example, if the user-oriented V L is the third listener orientation V 2,1is determined to be the closest, and the reduced transformation parameter W associated with this listener orientation pitch,2,1 is applied to the main rendered presentation R 1 In addition to the complete transformation parameter W associated with the second listener orientation V 2,0 note that the reduced transformation parameter W 2 may also be applied. Generally, when the main transformation parameter W pitch,2,1 and the reduced transformation parameter W associated with W 1 W 2 are used, the step of determining the modified transformation parameter W to apply to the main presentation pitch,2,1 involves determining the best match set of the main transformation parameters W L by (either selecting the closest set or by interpolation), and for the best match set of the main transformation parameters W 1 W 2 note that it involves determining the best match set of the reduced transformation parameters W 1 W 2 The modified transformation parameter W pitch,2,1 W pitch,2,-1 is obtained as a combination of these best match sets, which means that the output presentation is the best match main transformation parameters W L W 1 W 2 as well as the best match reduced transformation parameters W pitch,2,1 W pitch,2,-1 applied to the main presentation R 1 This is what is meant by being obtained by applying.

[0097] In FIG. 7, the multi-presentation encoder 13 and the interactive renderer 15 are shown alongside the parameter encoder 18 and the parameter decoder 19. The parameter encoder 18 may be part of the multi-presentation encoder 13 or may be provided externally. Similarly, the parameter decoder 19 may be part of the interactive encoder 15 or may be provided externally. As illustrated above, the multi-presentation encoder 15 is assumed to be implemented in a wearable device, such as a set of earphones or a set of headphones. Since the bandwidth for communicating with the wearable device may be limited (for example, the communication may be wireless such as Bluetooth (registered trademark)), it is beneficial if the amount of data transmitted to the interactive renderer 15 is limited. The amount of data, as explained above, may be reduced by using the down-conversion parameter W pitch for several listener positions. As will be explained, the amount of data transmitted to the interactive renderer 15 may alternatively or additionally be reduced by using efficient encoding and decoding of the conversion parameters W p , W pitch .

[0098] The parameter encoder 18 obtains the conversion parameters W 1 ,…,W P associated with the respective positions V 1 ,…,W P (optionally including one or more sets of down-conversion parameters W pitch ). The set of conversion parameters (and any down-conversion parameters) W 1 ,…,W P is combined into a vector W → . For time-frequency tile-based processing, the vector W → is generated for each time-frequency tile. For example, if three sets of conversion parameters W p are to be transmitted to the interactive renderer and each set of conversion parameters includes four complex values and one real value, the vector W →contains 27 vector elements. The main presentation R 1 Any default conversion parameter W 1 associated with can be included optionally. This is because they may already be stored in the interactive renderer 15.

[0099] For each frequency band b, a plurality of predetermined basis vectors u b,n can be used by the parameter encoder 18 to approximate the vector W → Overall, there are N b basis vectors for each frequency band, which means that the index is n = 1,..., N b It is assumed that the number N b of basis vectors may be the same across all frequency bands, or different numbers of basis vectors may be used for different frequency bands. For example, for lower frequency bands, fewer basis vectors can be used.

[0100] Using the basis vector u b,n the parameter encoder 18 determines a linear combination of basis vectors that forms a vector W → to approximate the vector W basis → Here, α

Number

[0101] The vector W → can exactly reconstruct the basis vector ub,n While it is possible to determine b,n , the parameter encoding performed by the parameter encoder 18 is generally not lossless. On the other hand, in some implementations, the residual vector R → is also determined by the parameter encoder 18

Number

[0102] The coefficient α n , and optionally the residual vector R, are included in the bitstream B transmitted from the parameter encoder 18 to the parameter decoder 19. The parameter decoder 19 stores a predetermined basis vector u b,n and uses the received coefficients and the above equation 22 to reconstruct the vector W basis → . Optionally, if the residual vector R → is also available, the parameter decoder 19 uses the above equation 23 to reconstruct the true transformation vector W → . The reconstructed vector W basis → represents the reconstruction of the transformation parameter ^W p , and the parameter processor 17 can use it to form the corrected reconstruction parameter W L . Alternatively, if W → can be reconstructed, W → represents the original transformation parameter W 1 , …, W P .

[0103] In some implementations, the conversion parameters W 1 ,…,W P and the orientations / positions V 1 ,…,V P associated therewith are transmitted together with the conversion parameters. For example, the data indicating the orientations / positions V 1 ,…,V P can be incorporated into the vector W → or encoded separately. For example, the conversion parameters W 1 ,…,W P and the determined rendering orientations / positions V 1 ,…,V P and their relative orientations / positions may vary over time. The positions V 1 ,…,V P are assumed to be pre-determined and may be available, for example, stored locally on the interactive renderer 15 as described above.

[0104] FIG. 8 shows a flowchart illustrating a method for processing audio according to some implementations. Further referring to FIG. 1, the method includes, in step S1, receiving, in a second processing module, at least one input signal. The first processing module may be a multi-presentation renderer 11 that receives one or more input audio signals from a database 10 having audio assets. The method then proceeds to step S2, which includes generating a main rendered presentation R 1 based on the at least one input audio signal and generating at least one additional presentation R 2 ,…,R P (P≧2). The main presentation R 1 and the at least one additional presentation R 2 ,…,R P are each associated with a different listener orientation V 2 ,…,V P .

[0105] In step S3, the conversion parameter W 2 , …, W P sets are determined for each additional presentation R 2 , …, R P . The conversion parameter W 2 , …, W P each set is configured to modify the main presentation R 1 into an additional presentation W p . Referring further to FIG. 9, step S3, which includes determining the conversion parameter, includes step S31 of determining the conversion parameter W p for an additional presentation, where the additional presentation is associated with a different yaw orientation compared to the main presentation, and step S32 of determining the reduced conversion parameter W pitch for a second additional presentation that has the same yaw orientation as the additional presentation but is associated with a different pitch orientation. That is, the method determines one or more sets of conversion parameters spanning a range of different main yaw orientations having a predetermined pitch, and then, for at least one of the main yaw orientations, determines one or more reduced sets of the conversion parameter W pitch describing how a presentation in the main yaw orientation having the predetermined pitch can be converted into a presentation having the same yaw orientation but a different pitch. The conversion parameters of the reduced set of conversion parameters indicate a real-valued gain for each channel, while the (non-reduced) set of conversion parameters indicates at least an N×N matrix, where N is the number of channels in the presentation.

[0106] Returning to FIG. 8, the method then proceeds to step S4, which includes transmitting the conversion parameter and the main presentation R 1 to a second processing module. The second processing module may implement the interactive renderer 15 as described above. The interactive renderer 15 obtains the user orientation / position V L or at least user orientation / position data indicating the orientation / position of the user's head.

[0107] In step S5, the interactive renderer 15 determines the orientation of the user and the orientation / position deviation value between the main presentation R 1 and the conversion parameter W 2 ,…,W P and the listener position V 1 ,…,V P associated therewith. Based on this orientation / position deviation value, in step 86, the interactive renderer 15 shifts the main presentation R 1 from the first listener orientation / position V 1 to the user orientation / position V L and determines the corrected conversion parameter W L . Determining the corrected conversion parameter W L may include selecting a set of conversion parameters W L associated with the orientation / position V p closest to the user orientation / position V p , or interpolating between at least two sets of conversion parameters Wp,…,W L associated with the listener orientation / position in the vicinity of the user orientation / position V P .

[0108] In step S7, the corrected conversion parameter W L is applied to the main presentation R 1 to form an output presentation.

[0109] To facilitate efficient low-bitrate communication between the first processing module and the second processing module, it is beneficial if the conversion parameters W 2 ,…,W P are encoded in an efficient manner. FIG. 10 is a flowchart showing a detailed embodiment of step S4 from FIG. 8 according to some implementations. The conversion parameters W 2 ,…,W P are combined into a vector W → . In step S41, the vector W → describing the conversion parameters is approximated by a linear combination of predetermined basis vectors u b,n . This approximation is the coefficient αn represented by a set, and the basis vector u b,n is used for the vector W → is sometimes referred to as the encoding of. The basis vector u b,n The number of is less than the number of elements of the vector W → This means that the scalar α n is W basis → is sometimes referred to as W → may represent an incomplete reconstruction of.

[0110] In step S42a, the scalar α n is sent to the second processing module. Optionally, in step S42b, W → and W basis → A residual vector R that describes the difference between what is called and → is determined and sent to the second processing module.

[0111] In the second processing module, the scalar α n (and optionally the residual vector R → ) is used in step S43 to reconstruct W b,n using the same predetermined basis vector u basis → (or if R → is available, then W → ). The process of reconstructing W → basis or W → is sometimes referred to as decoding the encoded conversion parameters.

[0112] Unless otherwise specified, as will be apparent from the following description, throughout this disclosure, descriptions using terms such as "processing", "calculating", "computing", "determining", "analyzing", etc. refer to actions and / or processes of a computer hardware or computing system, or similar electronic computing devices that manipulate and / or transform data represented as physical quantities, such as electronic quantities, into other data represented as physical quantities as well.

[0113] In the above description of exemplary embodiments of the present invention, it should be understood that various features of the present invention may be grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and assisting in the understanding of one or more of the various aspects of the invention. However, this method of disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, aspects of the invention lie in less features than all of the features of a single above-disclosed embodiment. Thus, the claims that follow the detailed description are expressly incorporated into this detailed description, and each claim stands on its own as a separate embodiment of the present invention. Further, some of the embodiments described herein include some features included in other embodiments but not others, and combinations of features of different embodiments are within the scope of the present invention and are intended to form different embodiments. This will be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments may be used in any combination.

[0114] Furthermore, some embodiments are described herein as a method or a combination of method elements that can be implemented by a processor of a computer system or by other means for performing functions. Thus, a processor having instructions for executing such a method or method elements forms means for executing the method or method elements. It should be noted that when a method includes several elements, for example, several steps, the order of such elements is not implied unless specifically stated. Further, the elements described herein in the device embodiments are an example of means for performing the functions executed by the elements for the purpose of implementing embodiments of the present invention. In the description provided herein, numerous specific details are set forth. However, it will be understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure the understanding of this description.

[0115] Thus, while specific embodiments of the present invention have been described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of the present invention, and it is intended to claim all such changes and modifications that fall within the scope of the present invention.

[0116] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE).

[0117] 〔EEE 1.1〕 A method for processing audio, comprising: in a first processing module, receiving an audio input including an audio channel, an object, metadata, or a combination thereof; in the first processing module, generating one or more rendered presentations and presentation conversion data of the audio input; in a second processing module, receiving user interactivity data and the one or more rendered presentations and presentation conversion data generated by the first process; and in the second processing module, generating an output presentation in response to the received user interactivity data, rendered presentation, and presentation conversion data.

[0118] 〔EEE 1.2〕 The output presentation is configured for headphone playback, as described in EEE 1.1.

[0119] 〔EEE 1.3〕 The user interactivity data indicates the orientation or position of the user's head, as described in EEE 1.1 or 1.2.

[0120] 〔EEE 1.4〕 The two processing modules are implemented on different devices having different processing capabilities and / or processing latencies, as described in any of the foregoing EEEs.

[0121] 〔EEE 1.5〕 The presentation conversion data represents a gain or input / output matrix having real or complex-valued coefficients, as described in any of the foregoing EEEs.

[0122] 〔EEE 1.6〕 The processing is applied as a function of time and frequency, as described in any of the foregoing EEEs.

[0123] 〔EEE 1.7〕 The first process is divided into two sub-processes, the first sub-process being a renderer that renders multiple presentations, and the second sub-process generating presentation conversion data, as described in EEE 1.1.

[0124] 〔EEE 1.8〕The second process includes a decorrelator stage, and the output of the decorrelator stage is mixed into the output presentation with a gain that depends on the presentation conversion data, according to any of the methods described in the foregoing EEE.

[0125] 〔EEE 1.9〕The user interactivity data includes the yaw and pitch angles (representation) of the user's head, and the presentation conversion data includes data elements for two or more yaw and / or pitch angles, according to any of the methods described in the foregoing EEE.

[0126] 〔EEE 1.10〕The method according to EEE 1.9, wherein the data elements for the yaw and pitch angles are represented individually as yaw contribution and pitch contribution.

[0127] 〔EEE 1.11〕The presentation conversion data is represented by a predetermined set of basis functions and a set of basis function weights, according to any of the methods described in the foregoing EEE.

Claims

1. A method of processing audio: In the first processing module, there is a step of receiving at least one input audio signal; The first processing module comprises the steps of generating a main rendered presentation and additional rendered presentations, each rendered presentation being associated with a first and a second listener orientation and / or position; The first processing module includes the step of determining transformation parameters for converting the main rendered presentation into the additional rendered presentation; The second processing module includes the steps of receiving the conversion parameters and the main rendered presentation generated by the first processing module; The second processing module includes the step of receiving user orientation and / or location data indicating the user's orientation and / or location; The second processing module includes the step of determining a deviation score based on the user's orientation and / or position and at least one of the first and second listener orientations and / or positions; The second processing module includes the step of determining a modified conversion parameter based on the conversion parameter and the deviation value; The second processing module includes the step of applying the modified transformation parameters to the main rendered presentation to generate an output presentation associated with the user's orientation and / or position, method.

2. A computer program product that, when the program is executed by a computer, causes the computer to perform the method described in claim 1.

3. A computer-readable storage medium storing the computer program described in claim 2.