Conversion of scene-based audio representations to object-based audio representations

By constructing a hybrid matrix and utilizing the generalized inverse of the object mapping matrix and scene mapping matrix, combined with the selection of amplitude preference coefficient, the conversion problem from SBA to OBA is solved, and a more discrete OBA rendering effect is achieved, and the reconstruction quality of audio scenes is improved.

CN120036013APending Publication Date: 2025-05-23DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072398.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-15
Filing Date
2023-09-25
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

现有技术难以有效将基于场景的音频格式(SBA)转换成基于对象的音频格式(OBA),尤其是难以在OBA格式中实现更离散的渲染效果。

Method used

By constructing a hybrid matrix suitable for converting the SBA input signal into the OBA output signal, the generalized inverse of the object mapping matrix and the scene mapping matrix are utilized, combined with the selection of amplitude preference coefficients, to achieve more discrete rendering of the OBA signal.

Benefits of technology

It realizes the effective conversion from SBA to OBA, providing a more discrete OBA rendering effect, making the reconstruction of audio scenes more realistic and flexible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036013A_ABST
    Figure CN120036013A_ABST
Patent Text Reader

Abstract

A hybrid matrix adapted to convert a scene-based audio (SBA) input signal into an object-based audio (OBA) signal is constructed such that the resulting OBA signal consists of object signals having amplitudes biased according to amplitude preference coefficients. The amplitude preference coefficients are selected to place dominant spatial audio objects in a smaller number of output object channels to provide a more discrete OBA rendering of the SBA input signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority from U.S. Provisional Application No. 63 / 379,081, filed on October 11, 2022, U.S. Provisional Application No. 63 / 479,236, filed on January 10, 2023, and U.S. Provisional Application No. 63 / 519,787, filed on August 15, 2023, the entire text of each of which is incorporated herein by reference. Technical Field

[0003] The present disclosure relates to representing acoustic scenes using multi-channel audio formats, and in particular, to conversion between different audio formats representing the same acoustic scene. Background Art

[0004] A set of audio signals may be processed and then transmitted through a transducer (e.g., a loudspeaker) with the goal of recreating a desired listening experience for one or more listeners. The set of audio signals may be referred to herein as a "multi-channel audio signal." The listening experience may be referred to herein as an "audio scene," and in particular, the term "target audio scene" refers to the desired listening experience (i.e., the listening experience that the multi-channel audio signal is intended to recreate).

[0005] A multi-channel audio signal will typically be associated with additional information that defines how a target audio scene relates to the multi-channel audio information. This additional information will include the name of the "format" of the multi-channel audio signal. As known in the art, typical formats include commonly known channel-based formats: stereo, 5.1, 7.1, etc. (collectively referred to as channel-based audio (CBA)). In the case of these CBA formats, the method of defining a target audio scene is based on each channel of the multi-channel audio signal being transmitted through a corresponding speaker, where the placement of the speakers around the listener is defined by the format. Typical formats also include object-based audio (OBA) formats, where a target audio scene is defined based on the transmission of each channel of the multi-channel audio signal to the listener, where the perceived DOA of each of the channels is defined by additional metadata known in the art. An example of an OBA format is Dolby Audio, developed by Dolby Laboratories of San Francisco, California, USA.

[0006] The audio channels associated with the time-varying DOA in the multi-channel signal may be referred to as dynamic objects, and the audio channels associated with the time-invariant DOA in the multi-channel signal may be referred to as static objects. By defining a static object for each of the channels in the original CBA format, the audio scene defined by the CBA format may be represented by the OBA format.

[0007] Typical formats also include scene-based audio (SBA) formats, in which a multichannel signal defines a target audio scene in terms of a target acoustic wavefield that should be reconstructed near the listening position. Scene-based formats do not specify the method by which the target acoustic wavefield should be generated. In addition, given the complexity of the acoustic wavefield, the multichannel audio signal may only attempt to define a subset of the information related to the acoustic wavefield. A common family of SBA formats is Ambisonics. The first-order Ambisonics (FOA) format defines a target audio scene by providing a multichannel audio file consisting of 4 channels, each of which defines a signal expected to be received by a corresponding ideal microphone positioned at a center point within the target acoustic wavefield, and each of which responds to the incident sound according to a specific directivity pattern.

[0008] According to the convention adopted in the field of Ambisonics production, the incident DOA of the sound is defined according to a 3-dimensional coordinate system, where the X-axis points forward, the Y-axis points to the left, and the Z-axis points upward. In the FOA format, the 4 microphone directivity patterns are chosen to be an omnidirectional pattern plus 3 dipole patterns, where the 3 dipole patterns are aligned with the X, Y, and Z axes respectively. For example, an ideal dipole microphone aligned with the X-axis will capture the incident sound with a gain equal to x when exposed to the incident sound wave from the direction defined by the unit vector (x,y,z). The ideal omnidirectional microphone pattern can be considered to have a received gain of 1, regardless of the incident direction of the sound wave. Summary of the invention

[0009] A mixing matrix adapted to convert a scene-based audio input signal into an object-based audio output signal is constructed such that the resulting object-based audio signal consists of object signals having amplitudes biased according to amplitude preference coefficients. The amplitude preference coefficients are selected to place dominant spatial audio objects in a smaller number of output object channels to provide a more discrete object-based rendering of the scene-based audio input signal.

[0010] In some embodiments, a method includes: determining an object mapping matrix that defines linear mixing characteristics that map audio objects from an object-based format to a scene-based format; determining a cost factor for each audio object of the object-based format; determining a scene mapping matrix as a generalized inverse of the object mapping matrix, wherein the scene mapping matrix is ​​determined so as to minimize a sum of weighted energies of the audio objects, wherein the weighted energy of each particular audio object is scaled according to its corresponding determined cost factor; and generating an object-based audio signal comprising an audio object signal that is a mixture of an audio signal from a scene-based input signal according to the scene mapping matrix.

[0011] In some embodiments, the scene-based input signal is an M-channel multi-channel audio signal, each cost factor varies according to the amplitude preference of its corresponding audio object, and the amplitude preference of each audio object is determined from a weighted sum of elements of the matrix C, where C is the M x M covariance of the M-channel scene-based input signal, and where weights are determined so as to form amplitude preference values ​​that approximate the object-based panning function.

[0012] In some embodiments, each of the audio objects is associated with an object position, the scene-based input signal is associated with a dominant direction, and each of the cost factors is defined to be lower for audio objects having an associated object position closer to the dominant direction.

[0013] In some embodiments, the audio object is a dynamic audio object having a position determined by video scene analysis.

[0014] In some embodiments, the method further comprises estimating the dominant direction and a directional bias coefficient from the scene-based input signal, the directional bias coefficient indicating a fraction of the scene-based input signal energy emanating from the dominant direction.

[0015] In some embodiments, each cost factor varies according to an amplitude preference of its corresponding audio object, and the amplitude preference varies according to an incident direction of the audio object, the dominant direction, and the direction deviation coefficient.

[0016] In some embodiments, the function provides a larger value of the amplitude preference when the incident direction is closer to the dominant direction.

[0017] In some embodiments, the scene-based input signal is an M-channel multi-channel audio signal, and the dominant direction V dom is maximized A unit vector of values ​​of , where C is the M x M covariance of the M-channel scene-based input signal, and where the “*” operator indicates the transpose.

[0018] In some embodiments, the dominant directions are formed by elements of the covariance matrix C.

[0019] In some embodiments, the scene-based input signal is defined according to a first-order Ambisonics translation function.

[0020] In some embodiments, the scene-based input signal is split into two or more sub-band scene-based signals according to a frequency selective filtering process, wherein for each sub-band, the respective scene-based sub-band signal is converted into a separate object-based sub-band signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Embodiments disclosed herein will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0022] Figure 1A A converter for converting an SBA representation to an OBA representation is described for use in accordance with one or more embodiments.

[0023] Figure 1B is a diagram of an arrangement of sound emitting objects around a central reference point in accordance with one or more embodiments.

[0024] Figure 2 is a diagram of a scene mapping matrix generator for determining a scene mapping matrix in accordance with one or more embodiments.

[0025] Figure 3 is a diagram illustrating determination of an object audio signal according to one or more embodiments.

[0026] Figure 4 is a diagram showing determination of an object audio signal including determining a location of an object according to one or more embodiments.

[0027] Figure 5 is a diagram showing determination of an object audio signal including determining object locations for a set of subbands according to one or more embodiments.

[0028] Figure 6 is a block diagram of a system for detecting dominant spatial objects to generate amplitude preference coefficients for an SBA mapping matrix according to one or more embodiments, wherein the SBA mapping matrix places the detected dominant spatial audio objects in a smaller number of output object channels in an OBA format, thereby providing a more discrete OBA rendering of an SBA input signal.

[0029] Figure 7 is a flow diagram of an example process for converting an SBA to an OBA representation in accordance with one or more embodiments.

[0030] Figure 8is a block diagram of an example hardware architecture suitable for implementing the systems and methods described with reference to FIGS. 1-7 . DETAILED DESCRIPTION

[0031] Techniques related to converting audio signals from one format to another are described herein. In the following description, for the purpose of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure, as defined by the claims, may include some or all of the features in these examples (alone or in combination with other features described below) and may further include modifications and equivalents of the features and concepts described herein.

[0032] In the following description, various methods, processes, and procedures are described in detail. Although certain steps may be described in a particular order, this order is primarily for convenience and clarity. Certain steps may be repeated more than once, may occur before or after other steps (even if those steps are otherwise described in another order), and may occur in parallel with other steps. It is only necessary for a second step to follow a first step if the first step must be completed before the second step can be started. When this is not clear from the context, this will be explicitly noted.

[0033] In this document, the terms "and", "or", and "and / or" are used. Such terms should be understood to have an inclusive meaning. For example, "A and B" may mean at least the following: "both A and B", "at least both A and B". As another example, "A or B" may mean at least the following: "at least A", "at least B", "both A and B", "at least both A and B". As another example, "A and / or B" may mean at least the following: "A and B", "A or B". When an exclusive-OR is desired, this will be specifically noted (e.g., "A or B", "at most one of A and B").

[0034] This document describes various processing functions associated with structures such as blocks, elements, components, circuits, etc. In general, these structures can be implemented by a processor that is controlled by one or more computer programs, such as those described in the reference numerals. Figure 8 describe.

[0035] Overview

[0036] The disclosed embodiments relate to a method for processing an SBA signal into an image that can be rendered by an OBA signal (e.g., Dolby Digital). ). This processing allows various playback and / or intermediate processing systems to flexibly utilize SBA signals in an OBA listening environment.

[0037] SBA is a three-dimensional (3D) audio format that allows accurate capture, efficient delivery, and rendering of 3D audio sound fields on any device, such as headphones, arbitrary speaker configurations, sound bars, etc. The SBA signal includes several channels that describe the audio scene from the listening position. An example SBA format is High-Order Ambisonics (HOA). Unlike the CBA format, the HOA transmission channel contains a speaker-independent sound field representation that can be decoded into the listener's speaker settings. SBA allows audio content producers to represent audio scenes based on source directions rather than speaker positions, and provides listener flexibility regarding speaker layout and the number of speakers used for playback.

[0038] OBA is a format that treats each sound source as an independent object with its own metadata (such as position, volume, and direction). OBA allows audio to be dynamically rendered based on the listener's speaker layout, the listener's position, and the acoustic properties of the listening environment.

[0039] Background and overview of Ambisonics

[0040] There are methods for mapping OBA to SBA. Such mapping can be generally achieved by translating the function P 1 (x, y, z), the translation function maps the DOA of the Ambisonics audio scene (in the form of a unit vector (x, y, z)) into a column vector of N gain values ​​corresponding to the gains used to render the N objects, respectively. In the example of N=4, it would correspond to the FOA, and translating this DOA format to N=4 objects would be conceptually based on Equation 1 shown below:

[0041]

[0042] In the Ambisonics format, there are a variety of conventions for defining gain scales and channel ordering of signals. The methods, devices, and systems described herein can be applied to Ambisonics signals that adhere to alternative scale and channel order conventions without loss of generality. While CBA formats with 7 or more channels are becoming more widespread, scene-based 4-channel formats (such as FOA) will attempt to define a target audio scene with a relatively small number (e.g., 4) of audio channels.

[0043] The audio scene can be represented in terms of a second-order or third-order Ambisonics format (called HOA) consisting of 9 or 16 channels respectively. The associated panning function P for converting this second-order or third-order HOA into an object 2 (x,y,z) and P 3 (x, y, z) can be implemented according to the principles described in the translation functions of Equations 2 and 3, respectively.

[0044]

[0045] SBAs such as HOA are useful because they allow complex audio scenes to be represented in multi-channel signals that allow for easy manipulation and analysis using off-the-shelf audio processing tools. However, these methods are limited to converting OBAs to SBAs. It is desirable but more difficult to convert scene-based formats to channel-based or object-based formats. The disclosed embodiments perform this mapping to provide SBA to OBA conversions.

[0046] Object-Based Audio Overview

[0047] The object-based audio signal includes an audio signal intended to be transmitted from a group of audio emitting devices, each of which is positioned at a specified position relative to a central listening position. The OBA format may be formed by N objects, where the value of N may be greater than or equal to 1. Such an object may be represented, for example, by an object n (n=1, 2, ..., N), which is associated with the audio signal O n (t) and the object incident direction vector V n =[x n ,y n ,z n ] T , where the object incident direction vector defines the spatial position of the audio object in the 3D audio scene.

[0048] Preferably, without loss of generality, the parameter V n is a unit vector such that: The audio signal O(t) can be represented by an [N×1] column vector, which is formed by N object audio signals, {O n (t): n = 1..N}. In an alternative preferred embodiment, the object incident direction can also vary according to time: V n (t).

[0049] Conversion of scene-based audio to object-based audio

[0050] Figure 1AA converter for converting an SBA representation to an OBA representation is described in accordance with one or more embodiments. An SBA signal source 101 generates an SBA signal, which is received by an SBA to OBA converter 102 in the form of a bitstream that may include metadata. The SBA signal is converted by the SBA to OBA converter 102 into an OBA signal. The OBA signal may be used by various downstream receiving devices 103, including but not limited to: mobile devices (e.g., smartphones, tablets, home entertainment systems, automotive infotainment systems, etc.). The OBA signal may be rendered for playback in an OBA format, a CBA format (e.g., 5.1, 7.1 surround sound), a two-channel format (e.g., for headphones, earbuds, etc.). Alternatively, the OBA signal may be further processed, transmitted to other devices, stored or

[0051] The disclosed embodiments may be implemented in an audio encoder, decoder, intermediate processing device, or general processing environment. In an encoder or audio content creator environment, a pre-processing module may implement the disclosed embodiments. For example, the disclosed embodiments may be used by an encoder or content creator to incorporate SBAs (e.g., FOA, HOA) into OBA productions (e.g., Dolby Digital Audio) by converting SBA channels into a set of objects using the disclosed embodiments. Production).

[0052] Alternatively or additionally, the disclosed embodiments may be implemented after a decoding process in a listening environment, where an SBA stream is produced at the output of the decoder, but audio objects are required for the purpose of rendering the audio in the listening environment. For example, if the SBA audio is to be rendered to a speaker in an OBA format, then the disclosed embodiments may be used to convert the channels of the SBA signal into an OBA signal, which is rendered into a speaker signal for speaker playback or rendered binaurally for playback on headphones / earbuds.

[0053] The processing blocks implementing the disclosed embodiments may be implemented in audio and / or audio-visual environments, such as mobile devices, wearable devices, home entertainment systems, smart phones, tablets, virtual reality (VR), augmented reality (AR) and mixed reality (MR) headsets / goggles / glasses, gaming consoles, automotive infotainment systems, and any other device capable of processing and / or rendering OBA.

[0054] Example SBA workflow

[0055] The example SBA workflow generally includes three stages: production, transmission and reproduction. In the production stage, the mixing engineer creates 3D audio content in HOA by mixing various audio sources (e.g., feeds from live microphones, stems, Ambisonics microphones, etc.) using appropriate tools to perform HOA transformations. In the transmission stage, the HOA signal group is compressed and sent to the end user in an audio bitstream (e.g., MPEG-H audio bitstream) or any other suitable bitstream format or transmission mechanism. In the reproduction stage, an audio decoder (e.g., MPEG-H audio decoder) at the user terminal receives and decodes the audio bitstream to retrieve the HOA signal. The HOA signal can then be further manipulated and customized (e.g., rotating the sound field in VR applications or performing audio "zooming" in the desired direction). Finally, the HOA renderer creates an appropriate feed for the reproduction device. In some embodiments, dialogue, commentary, or audio description can be sent as separate audio objects as needed.

[0056] In the reproduction stage, content producers can use the disclosed embodiments to convert SBA signals into OBA signals to perform specific tasks that cannot be performed using CBA or SBA. For example, in a live sports production scenario, a mixing engineer can deliver commentary in two different languages ​​(e.g., English, Spanish) as audio objects (separate from the mix), allowing English-speaking end users to select English commentary using their respective playback devices and Spanish-speaking end users to select Spanish commentary using their respective playback devices.

[0057] In some embodiments, the SBA signal may be analyzed and used to determine if any audio objects should be moved so that the objects are better positioned to 'align' with the dominant SBA channel. In some embodiments, an external position source may be used, such as video analysis.

[0058] Derivation of SBA to OBA mapping matrix

[0059] To convert SBA to OBA, the disclosed embodiments provide a generalized inverse of the OBA to SBA mapping. The derivation of the SBA to OBA mapping is now referred to Figure 1B Discuss.

[0060] Figure 1B An arrangement 100 of sound emitting devices 111, 112, 113 arranged around a central reference position 120 is shown according to one or more embodiments. The sound emitted by each emitting device 111, 112, 113 is incident at the central reference position 120 from an incident DOA 121, 122, 123, respectively. A 3-channel (N=3) OBA signal is obtained according to Figure 1B The arrangement 100 is rendered in FIG. 1 (t) can be transmitted from the transmitter 111, resulting from the unit vector V1 The specified direction 121 is incident on the sound wave at the central reference position 120. Similarly, the audio signal O 2 (t) and O 3 (t) can be transmitted from transmitters 112 and 113, respectively, from the unit vector V 2 and V 3 The designated directions 121 and 122 are incident on the central reference position 120 .

[0061] The audio scene around the center reference position 120 is composed of the object audio signal O 1 (t), O 2 (t) and O 3 (t) and the associated incident direction V 1 、V 2 and V 3 The audio scene can also be composed of M audio channels (S 1 (t), S 2 (t),…S M The signal vector S(t) represents an SBA signal vector of M SBA channels, wherein each audio signal is formed by the sum of the incident audio signals at the center reference position 120, and wherein each incident audio signal is scaled according to a panning gain function associated with the SBA channel. The signal vector S(t) represents an [M×1] column vector formed by M SBA signals, {S m (t):m=1..M}.

[0062] In an embodiment, the translation gain function maps each incident arrival direction into an [M×1] column vector of translation gain. These principles are exemplarily illustrated in conjunction with the following equation 4:

[0063]

[0064] Alternatively, for simplicity, the translation gain function can be expressed in terms of a [3×1] unit vector [x, y, z] T = V. The principle of the exemplary implementation in Equation 4 can be written in a more concise form as Equation 5:

[0065]

[0066] In an embodiment, the object mapping matrix E is determined according to the principle described in conjunction with equation 8, so that its generalized inverse can be determined according to the principle described in conjunction with equations 10 and 11. 1 (t), O 2 (t), ...O N (t) and the associated arrival direction vector V 1 、V 2 ,…V NThe composed object-based format can be converted into an M-channel scene-based format according to the principle of the exemplary implementation in Equation 6: 1 (t), S 2 (t),…S M (t) (defined by the scene-based translation function in Equation 5):

[0067]

[0068] The principle of the exemplary implementation in Equation 6 can be rewritten as shown in Equation 7:

[0069] S(t)=E×O(t). (7)

[0070] where S(t) and O(t) are scene-based and object-based signals, respectively, and the [M×N] object mapping matrix E is given by:

[0071]

[0072] The object-based signal vector can be determined by multiplying the scene mapping matrix D by the scene-based signal vector. These principles can be exemplarily illustrated by the following equation 9:

[0073] O(t)=D×S(t), (9)

[0074] where the [N×M] scene mapping matrix D is selected to satisfy Equation 10:

[0075] E×D=I M , (10)

[0076] And I M is the [M×M] identity matrix.

[0077] In an embodiment, because the number of objects N is greater than the number of scene-based channels M (such that N>M), there will typically be more than one scene mapping matrix D that satisfies Equation 10, and any such scene mapping matrix D that satisfies Equation 10 is called the generalized inverse of matrix E.

[0078] The expanded form of the scene mapping matrix D is shown according to the principle of Equation 11:

[0079]

[0080] The D matrix (shown as described below Figure 2The function of D in 205 is to define the manner in which the OBA signal is generated from the SBA signal (according to the principles discussed in conjunction with Equation 9), and in an embodiment, it is desirable to select D so as to increase or decrease the power in some OBA signals while D still satisfies Equation 10. In another embodiment, the allowed amplitude of each object channel (m=1, 2, ..., M) can be determined by the parameter a m definition.

[0081] Because there is more than one scene mapping matrix D that satisfies Equation 10, the scene mapping matrix D that satisfies Equation 10 can be associated with a cost function β(D), as shown in Equation 12:

[0082]

[0083] where the generalized inverse matrix D is chosen to minimize the cost function β(D). m A higher contribution of the power of row m of the matrix D to the cost function is associated.

[0084] According to alternative terminology, each object can be associated with a cost factor:

[0085]

[0086] where a m is the allowed amplitude of each object channel m, so that the channel with lower amplitude preference a m Object channels that favor which objects will be allocated the most energy are associated with larger cost factors.

[0087] To minimize the cost function according to the principle of the exemplary implementation in Equation 15, the amplitude preference matrix A is calculated according to Equation

[0088] The principle of the exemplary implementation in equation 14 is defined as:

[0089]

[0090] The scene mapping matrix D may be determined such that it satisfies Equation 10 while minimizing the cost function β(D) according to Equation 15, where the scene mapping matrix D is defined based on the object mapping matrix E and the amplitude preference matrix A.

[0091] D=A×A×E * ×(E×A×A×E * ) -1 , (15)

[0092] in() -1 The operator indicates the matrix inverse, and () *The operator indicates the matrix transpose (or, if the object mapping matrix E contains complex coefficients, the Hermitian transpose as known in the art).

[0093] The principle of Equation 15 can alternatively be expressed as Equation 16:

[0094] D = A × (E × A) + , (16)

[0095] As is known in the art, () + The pseudoinverse of the indicated operator.

[0096] System for generating scene mapping matrices

[0097] As previously discussed, the SBA signal can be mapped to the OBA signal by first computing an object mapping matrix E that maps the OBA signal to the SBA signal according to the principles discussed in conjunction with Equation 8 and then using the amplitude preference matrix A to determine the generalized inverse of the object mapping matrix E for weighting the object channels according to a desired preference (e.g., a preference for dominant directions of arrival). In some embodiments, the m object channels a m The magnitude of the preference is determined according to the principles of Equation 18 below. The coefficients of matrices E and D may be pre-computed and stored in memory (e.g., of a playback or intermediate device) or calculated in real time for streaming audio according to the principles of Equation 17 discussed below.

[0098] Figure 2 A scene mapping matrix generator 200 is shown according to one or more embodiments. The scene mapping matrix generator 200 includes a determine object mapping matrix process 201 and a determine scene mapping matrix process 202. M object positions 204V m is used by the object mapping matrix process 201 to generate the object mapping matrix E 205 according to the principles described in conjunction with Equation 8. The determine scene mapping matrix process 202 converts the amplitude preference coefficient 207a m Combined with the object mapping matrix E 205 to form a scene mapping matrix 206D according to the principle of Equation 15 or Equation 16. In an embodiment, the M object positions 204 may be provided in the metadata of the SBA bitstream (e.g., an MPEG-H bitstream) or provided by an external source such as a video analyzer, as shown in FIG. Figure 6 describe.

[0099] Derivation of the amplitude preference coefficient

[0100] In an embodiment, analysis of the SBA signal may be used to estimate the dominant DOA (unit vector V dom) and a directional bias coefficient 0≤b≤1, which indicates the fraction of energy in the audio signal estimated to be emitted from the dominant direction. For each target audio channel (m=1,2,…,M), the amplitude preference coefficient a m According to the object position V m , dominant direction V dom and directional deviation b, as shown in Equation 17:

[0101] a m =f(V m ,V dom ,b). (17)

[0102] In an embodiment, the function f() is selected so that when V m -V dom When a is smaller and b is larger, m Will be bigger.

[0103] In an embodiment, the function f() is determined according to the principle of Equation 18:

[0104]

[0105] in <V m ,V dom > is the unit vector V m and V dom The dot product of .

[0106] When the direction deviation b is large, the function in Equation 18 will be the incident direction V m Closer to the dominant audio direction V dom The object provides a magnitude preference a m Larger values ​​of , where the amplitude preference varies more between different object channels.

[0107] Figure 3 Display includes Figure 2 The arrangement 300 of the elements of the scene mapping matrix generator 200, wherein the scene mapping matrix generator 200 is selected from M amplitude preference coefficients 315, a m and M object positions 316, V m The scene mapping matrix 317, D, is generated. The arrangement 300 may be included in an implementation in an audio playback device and / or an intermediate processing device or in a device that processes or renders an object-based signal (eg, Dolby ) in an encoder or decoder in any other device or a general purpose processor.

[0108] The M-channel SBA signal 311 (e.g., HOA signal) S(t) may be received in a bitstream (e.g., MPEG-H bitstream). The M-channel SBA signal 311 is combined (e.g., multiplied) with the scene mapping matrix 317D by the mixer 301 to generate the N-channel OBA signal 312O(t) according to the principle of Equation 7.

[0109] In an embodiment, the SBA signal 311 may be provided by, for example, a broadcaster at a "live" event, such as a sporting event or concert. The SBA signal may be generated from audio signals captured by one or more HOA microphones positioned at the event. For example, for a basketball game, HOA microphones (e.g., field microphones, ambient microphones) may be placed at opposite ends of a basketball court and center court, as well as mounted to the ceiling. The broadcaster may use a mixing console with an HOA panner to mix the microphone signals with commentary (e.g., in different languages) and output an SBA signal. The HOA panner creates the SBA signal based on the audio input (e.g., audio objects and field microphones) and the properties of the sound source (e.g., the location and width of the sound source in 3D space).

[0110] The mixing engineer may also apply various spatial effects to the SBA signal (e.g., rotating the sound scene to align with the camera perspective, mirroring, distorting, zooming to a specific direction). In some embodiments, the mixing console may also output OBA signals and CBA signals to increase flexibility for different listening environments. For example, dialogue, commentary in multiple languages, or audio description may be sent as separate OBA signals as needed. In some embodiments, the SBA signal may be reproduced via the SBA through headphones to a binaural rendering module known in the art and, for example, paired with a head mounted display (HMD) in a virtual reality (VR) or augmented reality (AR) production to allow the 3D sound field to adapt to the user's head rotation in real time.

[0111] The advantage of using SBA signals is that SBA signals alleviate the problems that may occur with OBA signals due to limited transmission bandwidth, complexity constraints of consumer devices, and scene manipulation. The SBA format is speaker-independent and therefore allows SBA content to be rendered on any speaker layout. The SBA format also enables users to personalize and interact with immersive audio content.

[0112] Reference again Figure 3 , the SBA signal 311 is processed by the analysis scene-based signal process 302 to determine a dominant direction 313 corresponding to a characteristic of the SBA signal 311 over a period of time approximately t, V dom (t) and deviation 310, b(t). Determine amplitude preference process 303 takes M object positions 316 (given by object direction vector V m) and the dominant direction vector 313 (V dom (t)) and the bias 310 (b(t)) are combined to form M amplitude preference coefficients 315 according to the principle of equation 18, a m , which are also the coefficients of the amplitude preference matrix A used in Equation 16 to generate the scene mapping matrix D. The object positions 316 and the M amplitude preference coefficients 315 are input to the scene mapping matrix generator 200, which outputs the scene mapping matrix 317, which is multiplied with the SBA signal in the mixer 301 to generate the OBA signal, as previously described with reference to Figure 2 The OBA signal is then stored and / or transmitted (eg, via an MPEG-H bitstream) to various OBA devices for playback of an OBA representation or a CBA representation of the original audio signal or for further processing before transmission to other downstream devices.

[0113] In some embodiments, object location 316 is part of the streamed SBA metadata (e.g., an MPEG-H bitstream) or is provided by an external source such as a video scene analyzer, as shown in FIG. Figure 6 In some embodiments, object location 316 may be static or dynamic. Some examples of dynamic locations include, but are not limited to, locations generated by video analytics tracking (e.g., a basketball or soccer ball on a sports field, a referee, coach, and / or any area where significant movement is detected by video analytics (e.g., a fight on a hockey field)) or pre-set locations (e.g., the location of a tailgate on a basketball field where the camera (and associated HOA microphone) is fixed in position / orientation) or any other location information (e.g., location information manually set by a content creator).

[0114] Calculate the dominant direction and deviation based on the covariance of the SBA signal

[0115] In some embodiments, the [M×M] covariance of the SBA input signal over a period of time approximately t is formed according to the principle of Equation 19:

[0116]

[0117] Wherein the window function r(τ) has a maximum value at about τ=0, and thus the window function r(τ-t) has a maximum value at about τ=t, thereby ensuring that the covariance C(t) represents the temporal properties of the scene-based signal at about time t. The covariance C(t) may be pre-computed and stored in a memory of a playback or intermediate processing device, included in the bitstream metadata of the scene-based signal (e.g., MPEG-H bitstream metadata), or computed in real time.

[0118] It should be appreciated that when the audio signal is represented by discrete time samples, the integration operation according to the principle of Equation 19 may be replaced by a discrete summation operation, as is known in the art.

[0119] The deviation b can be determined according to the principle of equation 20:

[0120]

[0121] where operator is the square of the Frobenius norm of C (the sum of the squares of the magnitudes of the elements of C), and tr(C) is the trace of C (the sum of the diagonal terms).

[0122] In an embodiment, the dominant direction V dom can be determined to maximize The unit vector of the values ​​of .

[0123] In an embodiment, the SBA signal may be translated according to a first-order Ambisonics translation function P(x, y, x)=P 1 (x, y, z), where P 1 (x,y,z) is defined in Equation 21,

[0124]

[0125] And the dominant direction V dom can be formed from three elements from the covariance matrix,

[0126]

[0127] where Re() indicates the real part of the coefficient (which may be required when the covariance matrix contains complex values), and the subscript C m,1 Indicates the element at column 1, row m of matrix C.

[0128] Figure 4 Display includes Figure 3 The arrangement 400 of elements, wherein the addition is adapted to adopt a dominant direction 313, V dom (t) and the deviation 310, b(t) and generate a set of object positions 316 to determine the object position process 410. Figure 3 In some embodiments, the object location 316 is part of the metadata of the stream or provided by an external source such as a video scene analyzer, as shown in FIG. Figure 6 describe. Figure 4 An embodiment of the invention determines an object position 316 based on the dominant direction 313 and the deviation 310 and provides the determined object position 316 to determine the amplitude preference process 303 . Figure 4The processes shown in can be applied to SBA signals streamed and / or retrieved from storage media. These processes can be included in an encoder or decoder of any source, receiver, or intermediate device and used in any application that will benefit from converting an SBA representation into an OBA representation.

[0129] refer to Figure 4 In the box 410, the object position 316, {V n :n=1..N} may include a number K of fixed object locations and L of dynamic object locations, where L+K=N. In a preferred embodiment, K≥M, and the K fixed object locations 316 are selected such that the object locations {V n :n=L+1..N} are roughly evenly distributed around the listener.

[0130] In yet another embodiment, L=1, and the dynamic object V 1 According to the dominant direction V of the scene-based signal S(t) dom Therefore, the object position 316 can be determined according to the principles discussed in conjunction with Equation 23:

[0131]

[0132] One aspect of the present invention is to convert an SBA signal into an OBA signal according to the principles discussed in conjunction with Equation 9, wherein the scene mapping matrix D is adapted to vary over time according to the characteristics of the SBA signal. It is known in the art to convert an SBA signal into a less discrete OBA signal O′(t) according to the following implementation:

[0133] O′(t)=D fix ×S(t), (24)

[0134] The scene mapping matrix D fix is fixed. fix is called the passive decoding matrix, and the resulting object-based signal O′(t) is called the passively decoded OBA signal. One example of a fixed decoding matrix known in the art is formed from the pseudo-inverse of the object mapping matrix E according to:

[0135] D fix =E + (25)

[0136] In some embodiments, the amplitude preference coefficient a of each channel object audio channel (n=1..N) is n It may be determined from the amplitude or power of the corresponding channel of the passively decoded object-based signal.

[0137] In another preferred embodiment, the amplitude preference coefficient a at time t n Determined by:

[0138]

[0139] Where O′ n (t) refers to the nth channel of the passively decoded OBA signal, and the window function r(τ) has a maximum value at about τ=0, and thus the window function r(τ-t) has a maximum value at about τ=t, thereby ensuring that the amplitude preference coefficient a n Derived from the power of the nth channel of the passively decoded OBA signal at approximately time t.

[0140] In yet another embodiment, the covariance matrix determined according to the principles discussed in conjunction with Equation 19 may be used in conjunction with the principles discussed in conjunction with Equation 24 to determine a according to n :

[0141]

[0142] in refers to the nth element on the diagonal of the matrix C. Therefore, a set of amplitude preferences (a n :n=1..N) is formed by the diagonal of the matrix C:

[0143] In yet another embodiment, the method of Equation 27 can be rewritten as:

[0144]

[0145] where h n,m1,m2 It can be defined as follows:

[0146] h n,m1,m2 ={D fix} n,m1 {D fix} n,m2 , (29)

[0147] Or, where the matrix D fix In the case of complex elements:

[0148]

[0149] In another embodiment, the amplitude preference coefficient a n can be determined according to the principles discussed in conjunction with Equation 28, where the coefficient h is, as discussed below, n,m1,m2 Determined by substitution method.

[0150] It should be appreciated that where Equation 5 shows a translation function P(V) that defines a translation rule for an SBA signal format, a translation function P'(V) may define a translation rule for an OBA signal format.

[0151]

[0152] The translation gain (g′) defined by the translation function P′(V) j :j=1..J) can be used to determine the target OBA signal:

[0153] In an embodiment, the object-based translation function P'(V) is defined according to a Vector-Based Amplitude Translation (VBAP) method, as is known in the art.

[0154] For any original audio signal with an associated direction of arrival U′, the contribution of the original audio signal to the SBA signal will result in a covariance proportional to:

[0155]

[0156] It should also be appreciated that for an original audio signal with an associated arrival direction U′, the nth channel of the OBA signal will have g′ according to the object-based panning function of Equation 31 n The expected magnitude of (U′).

[0157] In an embodiment, the gain factor h n,m1,m2 It is determined such that for each n=1...N and for a series of unit vectors U':

[0158] g′ n (U′)≈∑ m1=1 ∑ m2=1 h n,m1,m2 (C U′ ) m1,m2 , (33)

[0159] Or, more specifically, the error:

[0160] eerr n (U′)=(g′ n (U′)-∑ m1=1 ∑ m2=1 h n,m1,m2 (C U′ ) m1,m2 ) 2 (34)

[0161] is minimized when averaging over a series of arrival directions U′. n,m1,m2 Selected to minimize:

[0162] ∫∫ U′∈S2 e rr n (U′)dU′, (35)

[0163] Here, group S2 refers to the group of (2-dimensional) unit vectors on the surface of the unit sphere.

[0164] In an embodiment, the scaling factor h n,m1,m2 (where n=1...N, m1=1...M and m2=1...M) is defined such that the amplitude preference coefficient a defined according to the principle of Equation 28 is n (where n=1 . . . N) is analogous to the panning gain according to g'(U') when the covariance C is associated with an audio scene with dominant sounds at the arrival direction U'.

[0165] In another embodiment, the scaling factor h n,m1,m2 (where n=1...N, m1=1...M and m2=1...M) is defined such that the amplitude preference coefficient a defined according to the principle of Equation 28 is n Each of (where n=1...N) is similar to the corresponding OBA channel O n The expected amplitude when the covariance C is associated with an audio scene with dominant sounds at the arrival direction V.

[0166] Sub-band scene-based signal embodiment.

[0167] In an alternative embodiment, the scene-based signal 311 may be split into two or more sub-bands according to a frequency selective filtering process. For each sub-band, the corresponding SBA sub-band signal may be selected according to the above method (e.g., according to Figure 3 The arrangement 300 in FIG. 3 is converted into an OBA subband signal.

[0168] Figure 5 An example arrangement 500 is shown in which an SBA signal 541 is processed by a filter bank analysis process 510 to generate a plurality of SBA sub-band signals, such as 521a ... 521n. For each sub-band SBA signal, such as 521a ... 521n, a corresponding processing block, such as 501a ... 501n processes a sub-band scene-based signal, such as 521a ... 521n, to form a corresponding sub-band object-based signal, such as 531a ... 531n. The sub-band object-based signals, such as 531a ... 531n, are combined by a sub-band synthesis process 520 to form an object-based signal 542.

[0169] Figure 5 Each processing block 501a...501n may be based on, for example, Figure 3 The method of the method shown in the arrangement 300 is implemented wherein Figure 3 The object position 316 is determined by the determine object position process 410 (in Figure 5 (in Chinese). Figure 5In an embodiment of arrangement 500 in FIG. 4 , each processing block, eg, 501a . . . 501n , determines sub-band state data ( 551a . . . 551n , respectively) that may be used by determine object position process 410 to assist in determining object position 316 .

[0170] The sub-band status data 551a ... 551n may include data indicating the loudness of the scene-based signal in the corresponding sub-band. The sub-band status data 551a ... 551n may also include the dominant direction 313, V in the corresponding sub-band. dom and deviation 310 data, such as Figure 3 Displayed in.

[0171] The determine object position process 410 may determine the position of one or more dynamic objects based on a set of dominant directions determined for each subband (by processing blocks, e.g., 501, 512). When only one (L=1) dynamic object is provided by the determine object position process 410, the dynamic object position may be determined as an average of the dominant directions determined for all subbands. In an embodiment, the dynamic object position may be determined as a weighted average of the dominant directions determined for all subbands based on a set of band weights. The band weights may vary such that for each subband, the band weight is larger when the loudness and / or deviation of that band is larger.

[0172] When two or more dynamic objects are determined by the determine object position process 410, the positions of the dynamic objects may be formed according to various methods known in the art. In an embodiment, a k-means clustering algorithm is used to determine two or more centroids from the dominant directions determined for all sub-bands. In another embodiment, a weighted k-means clustering algorithm may be applied, where for each sub-band, the band weight is larger when the loudness and / or deviation of the band is larger.

[0173] Combining Visual Object Tracking with Ambisonics Object Extraction

[0174] Figure 6 is a method for detecting dominant space objects to generate a scene mapping matrix D( Figure 4 and 5 Block diagram of system 600 for determining amplitude preference coefficients of scene mapping matrix D according to 317 in FIG. 317 , wherein the scene mapping matrix D places detected dominant spatial audio objects in a smaller number of output object channels in an OBA format, thereby providing a more discrete OBA rendering of an SBA input signal.

[0175] System 600 includes SBA source 101, object tracker 602, object selector 603, determined amplitude preference process 303, scene map generator 200 (see Figure 2), mixer 301 and OBA device 103. SBA source 101 may be, for example, a broadcaster of a sporting event. Object tracker 602 may be, for example, a video analyzer. Object selector 603 may be a process for selecting a dominant object or other object of interest from a plurality of objects (e.g., based on transient or other information). Determined amplitude preference process 303, scene map generator 200 and mixer 301 are as previously described with reference to Figures 2 to 3 The OBA device may be any downstream device that renders the OBA signal for playback or further processing, including but not limited to mobile devices, home entertainment devices, automotive infotainment devices, headphones, intermediate processing devices, etc.

[0176] In this example embodiment, the SBA source 101 provides a video stream and an SBA audio stream (e.g., using MPEG-H transport). The video stream is input into an object tracker 602, which detects objects and their corresponding positions across a sequence of video frames (e.g., using k-means). An object selector 603 selects one or more dominant objects or objects of interest to map into an OBA signal (e.g., based on transient analysis). The object positions are provided by the object tracker 602 to a determined amplitude preference process 303, which determines the amplitude preference coefficients of a scene mapping matrix D, as previously described with reference to FIGS. 1 to 4. The amplitude preference coefficients and object positions are input into a scene mapping generator 200, which generates a scene mapping matrix 206 / 317 (D matrix), as described with reference to FIGS. 1 to 4. The scene mapping matrix 206 / 317 is input to the mixer 301, which uses the scene mapping matrix D to convert the SBA signal into an OBA signal, where the OBA channels are weighted according to the amplitude preference coefficients so that the dominant spatial audio objects are placed in a smaller number of output OBA channels to provide a more discrete OBA rendering of the SBA input signal.

[0177] When dynamic objects are generated, the above conversion process needs to react quickly to ensure that any new transient sound wave elements are detected so that the dynamic objects can be moved to the correct position before the transient event. In some embodiments, this can be achieved by ensuring that the dynamic objects only move smoothly at a relatively slow speed so that the loudness / timbre changes are not so unstable. It doesn't matter even if one or more of the dynamic objects are in the wrong position in the audio object scene, because no sound is generated near those dynamic objects.

[0178] In some embodiments, object tracker 602 analyzes the video signal to identify the dominant sonic object of interest in the scene (e.g., the basketball in the previous example). Based on the video analysis, the dominant object position is likely to be 'seen' moving in a nice, continuous manner. This object position can then be used by object separator 602 to generate an object-based scene (e.g., an Atmos audio scene where the Atmos object is placed exactly where the video analysis determined it should be). In some embodiments, the video analysis suggests a neighborhood, and subsequent audio analysis moves the object (slowly) within that neighborhood.

[0179] In some embodiments, there may be multiple groups of static objects that may be selected based on one or more trigger conditions that may come from video and / or audio analysis or some other input source. In this embodiment, the multiple groups of static objects may change dynamically. For example, in a basketball game, there may be two groups of static objects: one group at each end of the court. One end of the court where the game is currently being played may have an active corresponding first group of static objects, and when the ball moves to the opposite end of the court, the first group of static objects dynamically switches to a second group of static objects at the opposite end of the court, as can be determined by video analysis. In some embodiments, the position of the static objects in a particular group may be used to determine an amplitude preference 303 to generate an amplitude preference 315, which ensures that a particular group of objects is included in a smaller number of output OBA channels to provide a more discrete OBA rendering of the SBA input signal.

[0180] In some embodiments, a video analyzer processes sequential video frames of an audio scene and outputs object movement between frames. Processing may include object tracking, filtering, and data association. Some examples of object tracking include, but are not limited to, kernel-based tracking (e.g., mean-shift tracking), iterative object localization based on maximization of a similarity measure (e.g., Bhattacharyya coefficients), or contour tracking that iteratively evolves an initial object contour by minimizing contour energy using gradient descent. Filtering and data association may include incorporating prior information about a scene or object, handling object dynamics, and evaluating different hypotheses. Some examples of filters include, but are not limited to, Kalman filters or specific filters.

[0181] In some embodiments, the tracked objects are processed to determine the dominant objects based on the audio associated with the tracked objects (e.g., transient analysis). A dominant direction vector and deviation may be determined for one or more dominant objects, which may be used to determine the amplitude preference coefficient 314 of the scene mapping matrix 206 / 317 to be applied to the SBA signal 311, as described above with reference to FIGS. 1 to 4.

[0182] Example process

[0183] Figure 7 is a flow chart of an example process 700 for converting scene-based audio into an object-based representation according to one or more embodiments. Process 700 may be performed using, for example, reference Figure 8 The electronic device architecture 800 described herein is implemented.

[0184] In some embodiments, a method includes: determining an object mapping matrix defining linear mixing characteristics that maps audio objects from an object-based format to a scene-based format (701); determining a cost factor for each audio object of the object-based format (702); determining a scene mapping matrix as a generalized inverse of the object mapping matrix (703), wherein the scene mapping matrix is ​​determined so as to minimize a sum of weighted energies of the audio objects, wherein the weighted energy of each particular audio object is scaled according to its respective determined cost factor; and generating an object-based audio signal including an audio object signal that is a mixture of an audio signal from a scene-based input signal according to the scene mapping matrix (704). Each of these steps is previously described above.

[0185] Instance Computing Device

[0186] Figure 8 A block diagram of an example computing apparatus 800 suitable for implementing example embodiments of the present disclosure is shown. Apparatus 800 includes, but is not limited to, server and client devices, as previously described with reference to FIGS.

[0187] As shown, the device 800 includes a central processing unit (CPU) 801 that can execute various processes according to a program stored in, for example, a read-only memory (ROM) 802 or a program loaded from, for example, a storage unit 808 to a random access memory (RAM) 803. In the RAM 803, data required when the CPU 801 executes various processes is also stored as needed. The CPU 801, ROM 802, RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0188] The following components are connected to the I / O interface 805: an input unit 806, which may include a keyboard, a mouse, or the like; an output unit 807, which may include a display (e.g., a liquid crystal display (LCD)) and one or more speakers; a storage unit 808, which includes a hard disk or another suitable storage device; and a communication unit 809, which includes a network interface card, such as a (e.g., wired or wireless) network card.

[0189] In some implementations, input unit 806 includes one or more microphones at different locations (depending on the host device), enabling audio signals to be captured in various formats (eg, mono, stereo, spatial, immersive, and other suitable formats).

[0190] In some implementations, the output unit 807 includes a system with various numbers of speakers. The output unit 807 (depending on the capabilities of the host device) can render the audio signal in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0191] In some embodiments, the communication unit 809 is configured to communicate with other devices (e.g., via a network). A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811 (e.g., a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or another suitable removable medium) is installed on the drive 810 so that a computer program read therefrom is installed into the storage unit 808 as needed. It should be understood by those skilled in the art that although the computing device 800 is described as including the above components, in actual applications, some of these components may be added, removed, and / or replaced, and all such modifications or substitutions are within the scope of the present disclosure.

[0192] According to an example embodiment of the present disclosure, the above process may be implemented as a computer software program or implemented on a computer-readable storage medium. For example, the embodiments of the present disclosure include a computer program product, including a computer program tangibly embodied on a machine-readable medium, a computer program including a program code for executing the method. In such embodiments, the computer program may be downloaded and installed from a network via the communication unit 809, and / or installed from a removable medium 811, such as Figure 8 Displayed in.

[0193] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits (such as control circuit systems), software, logic, or any combination thereof. For example, the units discussed above may be controlled by control circuit systems (such as Figure 8 The CPU 801 of the other components of the combination of the control circuit system can be executed, so the control circuit system can perform the actions described in the present disclosure. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, microprocessor or other computing device (such as a control circuit system). Although various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flow charts, or illustrated and described using some other diagrams, it should be understood that the blocks, devices, systems, techniques or methods described herein can be implemented in (as non-limiting examples) hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0194] In addition, the various blocks shown in the flowchart may be viewed as method steps and / or operations resulting from the operation of computer program code and / or a plurality of coupled logic circuit elements constructed to implement the associated functions. For example, embodiments of the present disclosure include computer program products, including computer programs tangibly embodied on machine-readable media, computer programs containing program code configured to implement the methods described above.

[0195] In the context of the present disclosure, a machine-readable storage medium may be any tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-temporary and may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include the following: an electrical connection with one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0196] The computer program code for implementing the disclosed method can be written in any combination of one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment with a control circuit system, so that the program code, when executed by the processor of the computer or other programmable data processing equipment, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server or distributed throughout one or more remote computers and / or servers.

[0197] Although this file contains many specific implementation details, these should not be understood as limitations on the scope of the content that can be claimed, but should be understood as descriptions of features that can be specific to a particular embodiment. Certain features described in the context of a single embodiment in this specification may also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. In addition, although the features may be described above as working in certain combinations and / or even initially required, one or more features from the required combination may be deleted from the combination in some cases, and the required combination may involve a change in a sub-combination or a sub-combination. The logical flow depicted in the figure does not require the specific order or sequential order shown to achieve the desired result. In addition, other steps may be provided from the described process, or steps may be eliminated from the described process, and other components may be added to or removed from the described system. Therefore, other embodiments are within the scope of the claims below.

Claims

1. A method wherein include: determining, with at least one processor, an object mapping matrix defining linear blending characteristics for mapping audio objects from an object-based format to a scene-based format; determining, with the at least one processor, a cost factor for each audio object of the object-based format; determining, with the at least one processor, a scene mapping matrix as a generalized inverse of the object mapping matrix, wherein the scene mapping matrix is ​​determined so as to minimize a sum of weighted energies of the audio objects, wherein the weighted energy of each particular audio object is scaled according to its corresponding determined cost factor; and An object-based audio signal comprising audio object signals that are a mixture of audio signals from a scene-based input signal is generated with the at least one processor according to the scene mapping matrix.

2. The method of claim 1 , wherein the scene-based input signal is an M-channel multi-channel audio signal, each cost factor varies according to an amplitude preference of its corresponding audio object, and the amplitude preference of each audio object is determined from a weighted sum of elements of the matrix C, wherein C is the M x M covariance of the M-channel scene-based input signal, and wherein weights are determined so as to form amplitude preference values ​​that approximate an object-based panning function.

3. A method according to claim 1 or 2, wherein each of the audio objects is associated with an object position, the scene-based input signal is associated with a dominant direction, and each of the cost factors is defined to be lower for audio objects having an associated object position closer to the dominant direction.

4. The method according to claim 3, further comprising: include: The dominant direction and a directional bias coefficient are estimated from the scene-based input signal, the directional bias coefficient indicating the fraction of scene-based input signal energy emanating from the dominant direction. 5 . The method according to claim 4 , wherein each cost factor varies according to an amplitude preference of its corresponding audio object, and the amplitude preference varies according to an incident direction of the audio object, the dominant direction, and the direction deviation coefficient. The method of claim 5 , wherein the function provides a larger value of the amplitude preference when the incident direction is closer to the dominant direction.

7. The method according to claim 6, wherein the scene-based input signal is an M-channel multi-channel audio signal, and the dominant direction V dom is maximized A unit vector of values ​​of , where C is the M x M covariance of the M-channel scene-based input signal, and where the “*” operator indicates the transpose.

8. The method according to claim 7, wherein the dominant direction is formed by elements of the covariance matrix C.

9. The method according to any of the preceding claims, wherein the audio object is a dynamic audio object having a position determined by video scene analysis.

10. The method according to any of the preceding claims, wherein the scene based input signal is defined according to a first order Ambisonics translation function.

11. A method according to any of the preceding claims, wherein the scene-based input signal is split into two or more sub-band scene-based signals according to a frequency selective filtering process, wherein for each sub-band, the corresponding scene-based sub-band signal is converted into a separate object-based sub-band signal.

12. A non-transitory computer-readable storage medium storing instructions that, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 11.

13. A computing device, include: at least one processor; and A memory storing instructions which, when executed by the at least one processor, cause the computing device to perform the method according to any one of claims 1 to 11.