Method, apparatus, and system for encoding and decoding directional sound sources

JP7905494B2Active Publication Date: 2026-08-14DOLBY LABORATORIES LICENSING CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2026-08-14

Smart Images

  • Figure 0007905494000030
    Figure 0007905494000030
  • Figure 0007905494000031
    Figure 0007905494000031
  • Figure 0007905494000032
    Figure 0007905494000032
Patent Text Reader

Abstract

To provide a method for decoding audio data that actualizes representation of a complicated radiation pattern and efficient encoding.SOLUTION: An audio encoding method has the processes of: receiving a mono (monophonic) audio signal corresponding to an audio object and representation of a radiation pattern corresponding to the audio object; encoding the mono audio signal; and encoding a source radiation pattern to determine radiation pattern meta data. The radiation pattern includes sound levels corresponding to a plurality of sampling times, a plurality of frequency bands, and a plurality of directions. Encoding the radiation pattern relates to determining spherical harmonic function conversion of the representation of the radiation pattern, compressing the spherical harmonic function conversion, and obtaining the encoded radiation meta data.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims the interests of U.S. Patent Application No. 62 / 658,067, filed on 16 August 2018; U.S. Patent Application No. 62 / 681,429, filed on 6 June 2018; and U.S. Patent Application No. 62 / 741,419, filed on 4 October 2018. The contents of these applications are incorporated herein by reference in their entirety.

[0002] Technical field This disclosure relates to encoding and decoding of directional sound sources and auditory scenes based on multiple dynamic and / or moving directional sound sources. [Background technology]

[0003] Real-world sound sources, whether natural or artificial (speakers, instruments, voices, mechanical devices), radiate sound in an anisotropic manner. Characterizing the radiation pattern (or "directivity") of a sound source can be crucial for proper rendering, especially in the context of interactive environments such as video games and virtual / augmented reality (VR / AR) applications. In these environments, users typically interact with directional audio objects by walking around them, thereby changing their auditory perspective on the generated sounds (also known as 6-degree-of-freedom [DoF] rendering). Users can also grab and dynamically rotate virtual objects, which also requires rendering of different directions in the radiation patterns of the corresponding sound source(s). In addition to a more realistic rendering of the direct propagation effect from source to listener, radiation characteristics also play a major role in higher-order acoustic coupling between the source and its environment (e.g., the virtual environment in a game), thus influencing reverberation (i.e., the back-and-forth waves as in an echo). As a result, such reverberations can influence other spatial cues, such as perceived distance.

[0004] Most audio game engines offer some way to represent and render directional sound sources, but they are generally limited to simple directional gains that rely on a simple first-order cosine function or "sound cone" (e.g., a power cosine function) and a simple high-frequency roll-off filter definition. These representations are insufficient to represent real-world radiation patterns and are not well-suited to simplified / combined representations of multiple directional sound sources. [Overview of the project] [Means for solving the problem]

[0005] Various audio processing methods are disclosed herein. Some such methods may involve encoding directional audio data. For example, some methods may involve receiving a mono audio signal corresponding to an audio object and a representation of a radiation pattern corresponding to the audio object. The radiation pattern may include, for example, multiple sample times, multiple frequency bands and multiple direction-corresponding sound levels. Some such methods may involve encoding the mono audio signal and encoding the source radiation pattern to determine radiation pattern metadata. The encoding of the radiation pattern may involve determining a spherical harmonic transform of the representation of the radiation pattern and compressing the spherical harmonic transform to obtain encoded radiation pattern metadata.

[0006] Some such methods may involve encoding multiple directional audio objects based on a cluster of audio objects. The radiation pattern may represent a centroid that reflects the average sound level value for each frequency band. In some such implementations, multiple directional audio objects are encoded as a single directional audio object with a directivity corresponding to a time-varying, energy-weighted average of the spherical harmonic coefficients of each audio object. The encoded radiation pattern metadata may indicate the position of the cluster of audio objects, which is the average of the positions of each audio object.

[0007] Several methods may involve encoding group metadata about the radiation pattern of a group of directional audio objects. In some examples, the source radiation pattern may be rescaled for the amplitude of the input radiation pattern in a certain direction for each frequency to determine the normalized radiation pattern. According to some implementations, compressing spherical harmonic transforms may involve singular value decomposition, principal component analysis, discrete cosine transform, data-independent bases, and / or eliminating spherical harmonic coefficients of spherical harmonic transforms above a threshold order of spherical harmonic coefficients.

[0008] Several alternative methods may involve decoding audio data. For example, some such methods may receive an encoded core audio signal, encoded radiation pattern metadata, and encoded audio object metadata, and may involve decoding the encoded core audio signal to determine the core audio signal. Some such methods may involve decoding the encoded radiation pattern metadata to determine the decoded radiation pattern, decoding the audio object metadata, and rendering the core audio signal based on the audio object metadata and the decoded radiation pattern.

[0009] In some cases, the audio object metadata may include at least one time-varying 3-degree-of-freedom (3DoF) or 6-degree-of-freedom (6DoF) source orientation information. The core audio signal may include multiple orientation objects based on a cluster of objects. The decoded radiation pattern may represent a centroid that reflects the mean value for each frequency band. In some cases, rendering may be based at least in part on applying subband gains based on the decoded radiation data to the decoded core audio signal. The encoded radiation pattern metadata may correspond to a time- and frequency-varying set of spherical harmonic coefficients.

[0010] According to some implementations, the encoded radiation pattern metadata may include audio object type metadata. The audio object type metadata may, for example, represent parametric directional pattern data. The parametric directional pattern data may include cosine functions, sine functions, and / or cardioid functions. In some examples, the audio object type metadata may represent database directional pattern data. Decoding the encoded radiation pattern metadata and determining the decoded radiation pattern may involve querying a directional data structure containing audio object types and corresponding directional pattern data. In some examples, the audio object type metadata may represent dynamic directional pattern data. Dynamic directional pattern data may correspond to a time- and frequency-varying set of spherical harmonic coefficients. Some methods may involve receiving dynamic directional pattern data before receiving the encoded core audio signal.

[0011] Some or all of the methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, random-access memory (RAM) devices, read-only memory (ROM) devices, and other memory devices such as those described herein. Thus, various innovative aspects of the subject matter described herein can be implemented on one or more non-temporary media storing software. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by, for example, one or more components of a control system such as those disclosed herein. The software may include, for example, instructions for executing one or more of the methods disclosed herein.

[0012] At least some aspects of this disclosure may be implemented through a device. For example, one or more devices may be configured to perform at least partially the methods disclosed herein. In some implementations, the device may include an interface system and a control system. The interface system may include one or more network interfaces, one or more interfaces between the control system and a memory system, one or more interfaces between the control system and another device, and / or one or more external device interfaces. The control system may include at least one of a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, or discrete hardware components. Thus, in some implementations, the control system may include one or more processors and one or more non-temporary storage media operationally coupled to the one or more processors.

[0013] In some such examples, a control system may be configured to receive audio data corresponding to at least one audio object via an interface system. In some examples, the audio data may include a monophonic audio signal, audio object position metadata, audio object size metadata, and rendering parameters. Some such methods may involve determining whether the rendering parameters indicate a position mode or a directional mode, and if it is determined that the rendering parameters indicate a directional mode, rendering the audio data for playback through at least one loudspeaker according to the directional pattern indicated by the position metadata and / or size metadata.

[0014] In some cases, rendering audio data may involve interpreting audio object position metadata as audio object orientation metadata. Audio object position metadata may include, for example, x,y,z coordinate data, spherical coordinate data, and / or cylindrical coordinate data. In some cases, audio object orientation metadata may include yaw, pitch, and roll data.

[0015] In some examples, rendering audio data may involve interpreting audio object size metadata as directional metadata corresponding to directional patterns. In some implementations, rendering audio data may involve querying a data structure containing multiple directional patterns and mapping position metadata and / or size metadata to one or more of the directional patterns. In some cases, a control system may be configured to receive the data structure via an interface system. In some examples, the data structure may be received prior to the audio data. In some implementations, the audio data may be received in Dolby Atmos format. Audio object position metadata may correspond, for example, to world coordinates or model coordinates.

[0016] Details of one or more implementations of the subject matter described herein are shown in the accompanying drawings and the following description. Other features, aspects, and advantages will be evident from this description, drawings, and claims. Note that the relative dimensions in the following drawings may not be drawn to scale. Similar reference numbers and symbols in different drawings generally indicate similar elements. [Brief explanation of the drawing]

[0017] [Figure 1A] This is a flowchart illustrating the blocks of an example audio encoding method.

[0018] [Figure 1B] Blocks of a process that can be implemented by an encoding system for dynamically encoding per-frame directivity information for a directional audio object according to an example are shown.

[0019] [Figure 1C] Blocks of a process that can be implemented by a decoding system according to an example are shown.

[0020] [Figure 2] Figures 2A and 2B represent the radiation patterns of an audio object in two different frequency bands.

[0021] [Figure 2C] A graph showing examples of a normalized radiation pattern and a non-normalized radiation pattern according to an example is shown.

[0022] [Figure 3] An example of a hierarchy including audio data and various types of metadata is shown.

[0023] [Figure 4] A flowchart showing blocks of an audio decoding method according to an example is shown.

[0024] [Figure 5A] A drum cymbal is depicted.

[0025] [Figure 5B] An example of a speaker system is shown.

[0026] [Figure 6] A flowchart showing blocks of an audio decoding method according to an example is shown.

[0027] [Figure 7] An example of encoding a plurality of audio objects is shown.

[0028] [Figure 8] This block diagram shows an example of a component of an apparatus that may be configured to perform at least some of the methods disclosed herein.

[0029] Similar reference numbers and symbols in various drawings indicate similar elements. [Modes for carrying out the invention]

[0030] Aspects of this disclosure relate to the representation and efficient encoding of complex radiation patterns. Some such implementations may include one or more of the following: 1. A representation of a general sound emission pattern as time- and frequency-dependent Nth-order coefficients (N≧1) of a real-valued spherical harmonics (SPH) decomposition. This representation can also be extended depending on the level of the reproduced audio signal. Unlike the case where the directional source signal itself is a PCM representation such as HOA, a mono-object signal can be encoded separately from its directional information, which is represented as a set of time-dependent scalar SPH coefficients in various subbands. 2. An efficient encoding method to reduce the bitrate required to represent this information. 3. A solution that dynamically combines emission patterns so that a scene composed of several emitting sound sources can be represented by an equivalent reduced number of sources during rendering, while maintaining its perceptual quality.

[0031] One aspect of this disclosure relates to representing a general radiation pattern in order to complement the metadata for each mono audio object with a set of time / frequency-dependent coefficients that represent the directivity of the mono audio object projected onto an N-th order spherical harmonic base (N≧1).

[0032] The primary radiation pattern can be represented by a set of four scalar gain coefficients for a predefined set of frequency bands (e.g., 1 / 3 octave). This set of frequency bands is also known as a bin or subband. The bin or subband may be determined based on a short-time Fourier transform (STFT) or a perceptual filter bank for a single data frame (e.g., 512 samples, as in Dolby Atmos). The resulting pattern can be rendered by evaluating a spherical harmonic decomposition in the desired direction around the object.

[0033] Generally, this radiation pattern is a characteristic of the source and may remain constant over time. However, updating this set of coefficients at regular time intervals can be beneficial to represent dynamic scenes where objects rotate or change, or to ensure that data is randomly accessible. In the context of dynamic auditory scenes with moving objects, the result of object rotation can be directly encoded in time-varying coefficients without requiring a separate, explicit encoding of object orientation.

[0034] Each type of sound source typically has a characteristic radiation / emission pattern that differs depending on the frequency band. For example, a violin may have a very different radiation pattern than a trumpet, drum, or bell. Furthermore, sound sources like musical instruments may radiate differently at pianissimo and fortissimo performance levels. As a result, the radiation pattern can be a function not only of the direction around the sound-producing object but also of the pressure level of the radiating audio signal, and the pressure level can also vary over time.

[0035] Therefore, instead of simply representing the sound field at a point in space, some implementations involve encoding audio data corresponding to the radiation pattern of an audio object so that it can be rendered from different vantage points. In some cases, the radiation pattern may be a radiation pattern that varies with time and frequency. In some cases, the audio data input to the encoding process may include multiple channels (e.g., 4, 6, 8, 20 or more channels) of audio data from directional microphones. Each channel may correspond to data from a microphone at a specific position in space around the sound source, from which the radiation pattern can be derived. Assuming that the relative direction from each microphone to the sound source is known, this can be achieved by numerically fitting a set of spherical harmonic coefficients such that the resulting spherical function best matches the observed energy levels in various subbands of each input microphone signal. For example, see the methods and systems described in relation to the international application PCT / US2017 / 053946, “Method, Systems and Apparatus for Determining Audio Representations,” by Nicolas Tsingos and Pradep Kumar Govindaraju, which is incorporated herein by reference. In other examples, the radiation pattern of an audio object may be determined by numerical simulation.

[0036] Instead of simply encoding audio data from a directional microphone at a sample level, some implementations involve encoding monophonic audio object signals along with corresponding radiation pattern metadata representing the radiation patterns for at least some of the encoded audio objects. In some implementations, the radiation pattern metadata may be represented as spherical harmonic data. Some such implementations may also involve a smoothing process and / or a compression / data reduction process.

[0037] Figure 1A is a flowchart showing the blocks of an example audio encoding method. Method 1 may be implemented by a control system including, for example, one or more processors and one or more non-temporary memory devices (such as control system 815, described later with reference to Figure 8). As with other disclosed methods, not all blocks of Method 1 are necessarily executed in the order shown in Figure 1A. Furthermore, alternative methods may include more or fewer blocks.

[0038] In this example, block 5 is involved in receiving a mono audio signal corresponding to an audio object, and also in receiving a representation of the radiation pattern corresponding to the audio object. According to this implementation, the radiation pattern includes sound levels corresponding to multiple sample times, multiple frequency bands, and multiple directions. According to this example, block 10 is involved in encoding the mono audio signal.

[0039] In the example shown in Figure 1A, block 15 is involved in encoding the source radiation pattern to determine the radiation pattern metadata. According to this implementation, encoding the representation of the radiation pattern involves determining the spherical harmonic transform of the representation of the radiation pattern, compressing the spherical harmonic transform, and obtaining the encoded radiation pattern metadata. In some implementations, the representation of the radiation pattern may be rescaled for each frequency with respect to the amplitude of the input radiation pattern in a certain direction in order to determine a normalized radiation pattern.

[0040] In some cases, compressing a spherical harmonic transformation may involve discarding some higher-order spherical harmonic coefficients. Some such examples involve removing spherical harmonic coefficients of a spherical harmonic transformation that are above a threshold order, for example, above order 3, above order 4, above order 5.

[0041] However, some implementations may involve alternative and / or additional compression methods. According to some such implementations, compressing spherical harmonic transforms may involve singular value decomposition, principal component analysis, discrete cosine transform, data-independent bases, and / or other methods.

[0042] According to some examples, Method 1 may also involve encoding multiple directional audio objects as a group or "cluster" of audio objects. Some implementations may involve encoding group metadata relating to the radiation pattern of a group of directional audio objects. In some cases, multiple directional audio objects may be encoded as a single directional audio object with a directivity corresponding to a time-varying, energy-weighted average of the spherical harmonic coefficients of each audio object. In some such examples, the encoded radiation pattern metadata may represent centroids corresponding to the average tone level values ​​for each frequency band. For example, the encoded radiation pattern metadata (or related metadata) may indicate the position of the cluster of audio objects, which is the average of the positions of each directional audio object in the cluster.

[0043] Figure 1B shows a block of a process that may be implemented by the encoding system 100 to dynamically encode frame-by-frame directional information for a directional audio object, as an example. This process may be implemented via a control system, such as the control system 815 described later with reference to Figure 8. The encoding system 100 may receive a mono audio signal 101 that can correspond to a mono object signal as discussed above. The mono audio signal 101 may be encoded in block 111 and provided to serialization block 112.

[0044] In block 102, static or time-varying directional energy samples at different sound levels in a set of frequency bands relative to a reference coordinate system can be processed. The reference coordinate system can be determined in some coordinate space, such as a model coordinate space or a world coordinate space.

[0045] In block 105, frequency-dependent rescaling of time-varying directional energy samples from block 102 may be performed. For example, frequency-dependent rescaling may be performed according to the example shown in Figures 2A-2C, as described below. Normalization may be based on amplitude rescaling, for example, for high frequencies with respect to low frequencies.

[0046] Frequency-dependent rescaling may be renormalized based on the assumed capture direction of the core audio. Such assumed capture direction of the core audio may represent the listening direction relative to the sound source. For example, this listening direction may be called the gaze direction, where the gaze direction may be a direction with respect to the coordinate system (e.g., forward or backward).

[0047] In block 106, the rescaled directional output of 105 may be projected onto a spherical harmonic basis, thereby giving the coefficients of the spherical harmonics.

[0048] In block 108, the spherical coefficients of block 106 are processed based on information from the instantaneous sound level 107 and / or rotation block 109. The instantaneous sound level 107 may be measured in a certain direction at a certain time. The information from rotation block 109 may indicate the (arbitrary) rotation of the time-varying source orientation 103. In one example, in block 109, the spherical coefficients may be adjusted to account for a correction of the time dependence of the source orientation to the originally recorded input data.

[0049] In block 108, target level determination may be further performed based on equalization determined with respect to the assumed capture direction of the core audio signal. Block 108 may output a set of rotated spherical coefficients equalized based on the target level determination.

[0050] In block 110, the encoding of the radiation pattern may be based on a projection of the spherical coefficients associated with the source radiation pattern onto a smaller subspace, resulting in encoded radiation pattern metadata. As shown in Figure 1A, in block 110, the SVD decomposition and compression algorithm may be performed on the spherical coefficients output by block 108. In one example, the SVD decomposition and compression algorithm in block 110 may be performed according to the principles described later in relation to equations 11-13.

[0051] Alternatively, block 110 can be used to represent spherical harmonics in a space that leads to irreversible compression.

number

[0052] Such a syntax may include different sets of coefficients for different pressure / intensity levels of the sound source. Alternatively, if directional information is available at different signal levels and the source level cannot be further determined at playback, a single set of coefficients may be dynamically generated. For example, such coefficients may be generated by interpolating between low-level and high-level coefficients based on the time-varying level of the object audio signal at encoding time.

[0053] The input radiation pattern for a mono audio object signal may be "normalized" with respect to a given direction, such as the primary response axis (which may be the direction in which it was recorded or the average of multiple recordings), and the encoded directivity and final rendering may need to be consistent with this "normalization." In one example, this normalization may be specified as metadata. Generally, it is desirable to encode a core audio signal that would convey a good representation of the object's timbre if directivity information were not applied.

[0054] Directional encoding One aspect of this disclosure is directed towards implementing an efficient encoding scheme for directional information, since the number of coefficients increases quadratically with respect to the order of the decomposition. An efficient encoding scheme for directional information may be implemented, for example, for the final transmission of an auditory scene to an end-rendering device over a limited-bandwidth network.

[0055] Assuming 16 bits are used to represent each coefficient, a representation of a fourth-order spherical harmonic function in a 1 / 3 octave bandwidth would require 25 × 31^=12 kbits per frame. Refreshing this information at 30 Hz would require a transmission bitrate of at least 400 kbps, which is more than what current object-based audio codecs currently require to transmit both audio and object metadata. In one example, the radiation pattern is: G(θ i ,φi ,ω) Equation (1) It may also be expressed by...

[0056] In equation (1), (θ i ,φ i ), i∈{1…P} represents discrete colatitude angles θ∈[0,π] and azimuthal angles φ∈[0,2π] with respect to the sound source, P represents the total number of discrete angles, and ω represents the spectral frequency. Figures 2A and 2B show the radiation patterns of an audio object in two different frequency bands. Figure 2A may, for example, show the radiation pattern of an audio object in the frequency band 100–300 Hz, and Figure 2B may, for example, show the radiation pattern of the same audio object in the frequency band 1 kHz–2 kHz. Since low frequencies tend to be relatively close to omnidirectional, the radiation pattern shown in Figure 2A is relatively closer to circular than the radiation pattern shown in Figure 2B. In Figure 2A, G(θ0,φ0,ω) represents the radiation pattern in the direction of the principal response axis 200, while G(θ1,φ1,ω) represents the radiation pattern in any direction 205.

[0057] In some examples, the radiation pattern may be captured and determined by multiple microphones physically placed around a sound source corresponding to an audio object, while in other examples, the radiation pattern may be determined through numerical simulation. In the multiple microphone example, the radiation pattern may vary over time, for example, to reflect a live recording. The radiation pattern can be captured at a variety of frequencies, including low frequencies (e.g., <100Hz), intermediate frequencies (100Hz < and >1kHz), and high frequencies (>10kHz). The radiation pattern is sometimes also known as the spatial representation.

[0058] In another example, the radiation pattern is a captured radiation pattern G(θ) at a certain frequency in a certain direction. i ,φ i ,ω) may be reflected in the standardization. For example:

number

[0059] In Equation (2), G(θ0, φ0, ω) represents the radiation pattern in the direction of the main response axis. Referring again to FIG. 2B, in one example, the radiation pattern G(θ i , φ i , ω) and the normalized radiation pattern H(θ i , φ i , ω) can be seen. FIG. 2C is a graph showing examples of the normalized radiation pattern and the non-normalized radiation pattern according to one example. In this example, the normalized radiation pattern in the direction of the main response axis, represented as H(θ0, φ0, ω) in FIG. 2C, has substantially the same amplitude over the illustrated range of the frequency band. In this example, the normalized radiation pattern in the direction 205 (shown in FIG. 2A), represented as H(θ1, φ1, ω) in FIG. 2C, has a relatively higher amplitude at higher frequencies than the non-normalized radiation pattern represented as G(θ1, φ1, ω) in FIG. 2C. For a given frequency band, the radiation pattern may be assumed to be constant for notational convenience, but in practice, it may change over time, for example, by different bowing techniques used in stringed instruments.

[0060] The radiation pattern, or its parametric representation, may be transmitted. Preprocessing of the radiation pattern may be performed prior to its transmission. In one example, the radiation pattern or parametric representation may be preprocessed by a computational algorithm, an example of which is shown in relation to FIG. 1A. After preprocessing, the radiation pattern may be, for example

Number

[0061] In Equation (3), H(θ i , φ i , ω) represents the spatial representation,

Number

number

number

[0062] In equation (4), P n m (x) is a associated Legendre polynomial, of order m ∈ {-N…N}, and degree n ∈ {0…N}.

number

[0063] Other spherical bases may be used. Any method for performing spherical harmonic transformations on discrete data can be used. For example, matrix transformation

number

number

[0064] In equation (7),

number

[0065] pseudo inverse matrix

number

number

[0066] The regularized solution may also be applicable when the distribution of spherical samples contains a large amount of missing data. Missing data may correspond to regions or directions where directional samples are unavailable (for example, due to uneven microphone coverage). In many cases, the spatial distribution of samples is sufficiently uniform, and the identity weighting matrix W yields acceptable results. Also, P≫(N+1) 2 It is assumed that this is the case, and the spherical harmonic representation

number

number

[0067] Here, discrete frequency band ω k Let's consider k ∈ {1…K}. We can stack matrices H(ω) so that each frequency range is represented by a column of matrices.

number

[0068] In other words, the spatial representation H(ω) can be determined based on frequency bins / bandwidths / sets. As a result, the spherical harmonic representation is:

number

[0069] In equation (10),

number

number

number

number

[0070] Some embodiments may involve performing singular value decomposition (SVD), where,

number

number

number

[0071] O=(N+1) 2 In some cases, to achieve compression, the encoder is used.

number

[0072] In equation (12),

number

number

number

number

[0073] The following are three examples for transmitting truncated decomposition vectors and truncated right singular vectors: 1. The transmitter may independently transmit the encoded radiation T and the truncated right singular vector V' for each object. 2. Objects may be grouped, for example, according to a similarity index, and U and V may be computed as representative basis for multiple objects. Thus, encoded radiation can be transmitted per object, and U and V can be transmitted per group of objects. 3. The left and right singular matrices U and V may be pre-calculated on a large database of representative data (e.g., training data), and information about V may be stored on the receiver side. In some such examples, only encoded radiation may be transmitted per object. DCT is another example of a basis that may be stored on the receiver side.

[0074] Space encoding of directional objects When complex auditory scenes containing multiple objects are encoded and transmitted, spatial encoding techniques can be applied in a way that best preserves the auditory perception of the scene, replacing individual objects with fewer representative clusters. Generally, replacing groups of sound sources with representative "centroids" requires calculating aggregated / average values ​​for each metadata field. For example, the position of a cluster of sound sources can be the average of the positions of each sound source. By representing the radiation pattern of each source using spherical harmonic decomposition as described above (see, for example, equations 1-12), it is possible to linearly combine the sets of coefficients in each subband for each source to construct the mean radiation pattern for the cluster of sources. By calculating a time-varying, perceptually optimized representation that better preserves the original scene, it is possible to construct a time-varying, perceptually optimized representation that is weighted over time by loudness or the energy of the spherical harmonic coefficients.

[0075] Figure 1C shows a block of a process that may be implemented by an example decoding system. The block shown in Figure 1C may be implemented by a control system for a decoding device (such as control system 815, described later with reference to Figure 8), which includes, for example, one or more processors and one or more non-temporary memory devices. In block 150, metadata and an encoded core mono audio signal may be received and deserialized. The deserialized information may include object metadata 151, an encoded core audio signal, and encoded spherical coefficients. In block 152, the encoded core audio signal may be decoded. In block 153, the encoded spherical coefficients may be decoded. The encoded radiation pattern information may include an encoded radiation pattern T and / or matrix V. Matrix V is,

number

[0076] Object metadata 151 may include information about the relative direction from the source to the listener. For example, metadata 151 may include information about the distance and direction of the listener and the distance and direction of one or more objects in 6DoF space. For example, metadata 151 may include information about the relative rotation, distance, and direction of the source in 6DoF space. In the example of multiple objects in a cluster, the metadata field may reflect information about a representative "centroid" that reflects the aggregated / average value of the cluster of objects.

[0077] Next, the renderer 154 may render the decoded core audio signal and the decoded spherical harmonic coefficients. In one example, the renderer 154 may render the decoded core audio signal and the decoded spherical harmonic coefficients based on object metadata 151. The renderer 154 may determine the subband gain for the spherical coefficients of the radiation pattern based on information from the metadata 151, for example, the relative direction from the source to the listener. The renderer 154 may then render the core audio object signal based on the determined subband gain for the corresponding decoded radiation pattern(s), source and / or listener pose information (e.g., x, y, z, yaw, pitch, roll) 155. The listener pose information may correspond to the user's position and viewing direction in 6DoF space. The listener pose information may be received from a source local to the VR playback system, such as an optical tracking device. The source pose information corresponds to the position and orientation in space of the sound-emitting object. This can also be inferred from local tracking systems, for example, when a user's hands are tracked and they interactively manipulate an object that produces a virtual sound, or when a tracked physical prop / proxy object is used.

[0078] Figure 3 shows an example of a hierarchy containing audio data and various types of metadata. As with other drawings provided herein, the number and types of audio data and metadata shown in Figure 3 are provided for illustrative purposes only. Some encoders may provide the complete set of audio data and metadata shown in Figure 3 (dataset 345), while other encoders may provide only a portion of the metadata shown in Figure 3, for example, only dataset 315, only dataset 325, or only dataset 335.

[0079] In this example, the audio data includes a monophonic audio signal 301. The monophonic audio signal 301 is an example of what is sometimes referred to herein as a “core audio signal,” but in some examples, the core audio signal may include audio signals corresponding to multiple audio objects contained in the cluster.

[0080] In this example, the audio object location metadata 305 is expressed in Cartesian coordinates. However, in an alternative example, the audio object location metadata 305 may be expressed via other types of coordinates, such as spherical or polar coordinates. Thus, the audio object location metadata 305 may include 3-degree-of-freedom (3DoF) location information. According to this example, the audio object metadata includes the audio object size metadata 310. In an alternative example, the audio object metadata may include one or more other types of audio object metadata.

[0081] In this implementation, dataset 315 includes a monophonic audio signal 301, audio object position metadata 305, and audio object size metadata 310. Dataset 315 may be provided, for example, in the Dolby Atmos® audio data format.

[0082] In this example, dataset 315 also includes an optional rendering parameter R. According to some disclosed implementations, the optional rendering parameter R can indicate whether at least a portion of the audio object metadata of dataset 315 should be interpreted in its “normal” sense (for example, as position or size metadata) or as directional metadata. In some disclosed implementations, the “normal” mode may be referred to herein as the “position mode,” and the alternative mode may be referred to herein as the “directional mode.” Several examples are described below with reference to Figures 5A–6.

[0083] In this example, orientation metadata 320 contains angular information to represent the yaw, pitch, and roll of the audio object. In this example, orientation metadata 320 represents yaw, pitch, and roll as Φ, Θ, and Ψ. Dataset 325 contains enough information to orient the audio object for a 6-degree-of-freedom (6 DoF) application.

[0084] In this example, dataset 335 includes audio object type metadata 330. In some implementations, audio object type metadata 330 may be used to indicate the corresponding radiant pattern metadata. The encoded radiant pattern metadata may be used (for example, by a decoder or a device receiving audio data from a decoder) to determine the decoded radiant pattern. In some examples, audio object type metadata 330 may essentially indicate "I am a trumpet," "I am a violin," etc. In some examples, the decoding device may have access to a database of audio object types and corresponding directional patterns. According to some examples, the database may be provided with the encoded audio data or before the transmission of the audio data. Such audio object type metadata 330 may be referred to in this paper as "data directional pattern data."

[0085] In some examples, audio object metadata may represent parametric directional pattern data. In some examples, audio object metadata 330 may represent a directional pattern corresponding to a cosine function of a specified power, such as a cardioid function.

[0086] In some examples, the audio object metadata 330 may indicate that the radiation pattern corresponds to a set of spherical harmonic coefficients. For example, the audio object metadata 330 may indicate that the spherical harmonic coefficients 340 are provided in the dataset 345. In some such examples, the spherical harmonic coefficients 340 may be a set of spherical harmonic coefficients that vary with time and / or frequency, as described above, for example. Such information may require the largest amount of data compared to the rest of the metadata hierarchy shown in Figure 3. Therefore, in some such examples, the spherical harmonic coefficients 340 may be provided separately from the monophonic audio signal 301 and the corresponding audio object metadata. For example, the spherical harmonic coefficients 340 may be provided at the start of audio data transmission, before real-time operation (e.g., real-time rendering operation such as a game, movie, or music performance) begins.

[0087] According to some implementations, a decoder-side device, such as a device that provides audio to a playback system, may determine the capabilities of the playback system and provide directional information according to those capabilities. For example, even if the entire dataset 345 is provided to the decoder, only the usable portion of the directional information may be provided to the playback system in some such implementations. In some examples, the decoding device may decide which type (single or multiple) of directional information to use according to the capabilities of the decoding device.

[0088] Figure 4 is a flowchart showing the blocks of an example audio decoding method. Method 400 may be implemented by a control system for a decoding device (such as control system 815, described later with reference to Figure 8), which includes, for example, one or more processors and one or more non-temporary memory devices. As with other disclosed methods, not all blocks of Method 400 are necessarily executed in the order shown in Figure 4. Furthermore, alternative methods may include more or fewer blocks.

[0089] In this example, block 405 is involved in receiving an encoded core audio signal, encoded radiation pattern metadata, and encoded audio object metadata. The encoded radiation pattern metadata may include audio object type metadata. The encoded core audio signal may include, for example, a monophonic audio signal. In some examples, the audio object metadata may include 3DoF position information, 6DoF position and source orientation information, audio object size metadata, etc. In some cases, the audio object metadata may change over time.

[0090] In this example, block 410 includes decoding the encoded core audio signal to determine the core audio signal. Here, block 415 includes decoding the encoded radiation pattern metadata to determine the decoded radiation pattern. In this example, block 420 involves decoding at least a portion of other encoded audio object metadata. Here, block 430 involves rendering the core audio signal based on the audio object metadata (e.g., audio object position, orientation, and / or size metadata) and the decoded radiation pattern.

[0091] Block 415 may involve the behavior of various types, depending on the specific implementation. In some cases, audio object type metadata may indicate database directional pattern data. Decoding encoded radiant pattern metadata to determine the decoded radiant pattern may involve querying a directional data structure containing the audio object type and the corresponding directional pattern data. In some cases, audio object type metadata may indicate parametric directional pattern data, such as directional pattern data corresponding to a cosine function, a sine function, or a cardioid function.

[0092] According to some implementations, audio object metadata may represent dynamic directional pattern data, such as a set of spherical harmonic coefficients that vary with time and / or frequency. Some such implementations may involve receiving the dynamic directional pattern data before receiving the encoded core audio signal.

[0093] In some cases, the core audio signal received in block 405 may include audio signals corresponding to multiple audio objects contained in a cluster. According to some such examples, the core audio signal may be based on a cluster of audio objects that may contain multiple directional audio objects. The decoded radiation pattern determined in block 415 may correspond to the centroid of the cluster and may represent the average value for each frequency band of each directional audio object. The rendering process in block 430 may involve applying a subband gain to the decoded core audio signal that is at least partially based on the decoded radiation data. In some examples, after decoding the core audio signal and applying decoded directional processing to it, the signal may further be virtualized to its intended position relative to the listener position. This may involve using audio object position metadata and known rendering processes such as binaural rendering through headphones or rendering using loudspeakers in the playback environment.

[0094] As described above with reference to Figure 3, in some implementations, audio data may be accompanied by rendering parameters (shown as R in Figure 3). Rendering parameters may indicate whether at least some audio object metadata, such as Dolby Atmos metadata, should be interpreted in the usual way (e.g., as position or size metadata) or as directional metadata. The usual mode may be referred to as “position mode,” and the alternative mode may be referred to herein as “directional mode.” Thus, in some examples, rendering parameters may indicate whether at least some audio object metadata should be interpreted as direction relative to the speaker or as position relative to the room or other playback environment. Such implementations may be particularly useful for directional rendering using smart speakers with multiple drivers, as described below.

[0095] Figure 5A depicts a drum cymbal. In this example, the drum cymbal 505 is shown emitting sound with a directional pattern 510 having a substantially vertical primary response axis 515. The directional pattern 510 itself is also primarily vertical, with some spread from the primary response axis 515.

[0096] Figure 5B shows an example of a speaker system. In this example, the speaker system 525 includes multiple speakers / transducers configured to radiate sound in various directions, including upwards. The top speaker can be used in a conventional Dolby Atmos manner ("position mode") to render its position so that sound reflects off the ceiling, for example, to simulate a height / ceiling speaker (z=1). In some such cases, the corresponding Dolby Atmos rendering may include additional height virtualization processing to enhance the perception of audio objects having a specific position.

[0097] In other use cases, the same upward-emitting speaker(s))(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s))(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s))(s)(s)(s)(s)(s)(s)(s)(s)(s)(s)(s

[0098] Figure 6 is a flowchart showing the blocks of an example audio decoder method. Method 600 may be implemented by a control system for a decoder including, for example, one or more processors and one or more non-temporary memory devices (such as control system 815, described later with reference to Figure 8). As with other disclosed methods, not all blocks of Method 600 are necessarily executed in the order shown in Figure 6. Furthermore, alternative methods may include more or fewer blocks.

[0099] In this example, block 605 is involved in receiving audio data corresponding to at least one audio object. This audio data includes a monophonic audio signal, audio object position metadata, audio object size metadata, and rendering parameters. In this implementation, block 605 is involved in receiving this data via a decoding device interface system (such as interface system 810 in Figure 8). In some cases, the audio data may be received in Dolby Atmos® format. The audio object position metadata may correspond to world coordinates or model coordinates, depending on the specific implementation.

[0100] In this example, block 610 is involved in determining whether the rendering parameter indicates a position mode or a directional mode. In the example shown in Figure 6, if it is determined that the rendering parameter indicates a directional mode, block 615 renders the audio data for playback (e.g., playback through at least one loudspeaker, headphones, etc.) according to a directional pattern indicated by at least one of the position metadata or size metadata. For example, the directional pattern may be similar to that shown in Figure 5A.

[0101] In some examples, rendering audio data may involve interpreting audio object position metadata as audio object orientation metadata. Audio object position metadata may be Cartesian / x,y,z coordinate data, spherical coordinate data, or cylindrical coordinate data. Audio object orientation metadata may be yaw, pitch, and roll metadata.

[0102] According to some implementations, rendering audio data may involve interpreting audio object size metadata as directional metadata corresponding to directional patterns. In some such examples, rendering audio data may involve querying a data structure containing multiple directional patterns and mapping at least one of the position metadata or size metadata to one or more of the directional patterns. Some such implementations may involve receiving the data structure via an interface system. According to some such implementations, the data structure may be received before the audio data.

[0103] Figure 7 shows an example of encoding multiple audio objects. In this example, information 701, 702, 703, etc., for objects 1 to n may be encoded. In this example, in block 710, representative clusters for audio objects 701 to 703 may be determined. In this example, groups of sound sources may be aggregated and represented by representative "centroids." This involves calculating aggregated / average values ​​for metadata fields. For example, the position of a cluster of sound sources can be the average of the positions of each sound source. In block 720, radiation patterns for representative clusters can be encoded. In some examples, radiation patterns for clusters may be encoded according to the principles described above, with reference to Figure 1A or Figure 1B.

[0104] Figure 8 is a block diagram showing examples of components of an apparatus that may be configured to perform at least some of the methods disclosed herein. For example, apparatus 805 may be configured to perform one or more of the methods described above with reference to Figures 1A-1C, Figure 4, Figure 6 and / or Figure 7. In some examples, apparatus 805 may be a personal computer, a desktop computer, or other local device configured to provide audio processing, or may include them. In some examples, apparatus 805 may be a server, or may include a server. According to some examples, apparatus 805 may be a client device configured to communicate with the server via a network interface. The components of apparatus 805 may be implemented via hardware, via software stored on a non-temporary medium, via firmware, and / or a combination thereof. The types and numbers of components shown in Figure 8 and other figures disclosed herein are for illustrative purposes only. Alternative implementations may include more, fewer, and / or different components.

[0105] In this example, the device 805 includes an interface system 810 and a control system 815. The interface system 810 may include one or more network interfaces, one or more interfaces between the control system 815 and the memory system, and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). In some implementations, the interface system 810 may include a user interface system. The user interface system may be configured to receive input from the user. In some implementations, the user interface system may be configured to provide feedback to the user. For example, the user interface system may include one or more displays having corresponding touch and / or gesture detection systems. In some examples, the user interface system may include one or more microphones and / or speakers. According to some examples, the user interface system may include devices that provide haptic feedback, such as motors, vibrators, etc. The control system 815 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASICS), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic, and / or discrete hardware components.

[0106] In some examples, device 805 may be implemented as a single device. However, in some implementations, device 805 may be implemented as multiple devices. In some such implementations, the functionality of the control system 815 may be contained within multiple devices. In some examples, device 805 may be a component of another device.

[0107] Various exemplary embodiments of this disclosure may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. Generally, this disclosure is understood to also encompass devices suitable for performing the methods described above. For example, a device (spatial renderer) having memory and a processor coupled to the memory, wherein the processor is configured to execute instructions and perform the methods according to embodiments of this disclosure.

[0108] While various aspects of the exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or any other pictorial representation, it will be understood that any blocks, apparatus, systems, techniques, or methods described herein may be implemented, in non-limiting examples, in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers, or other computing devices, or any combination thereof.

[0109] Furthermore, the various blocks shown in the flowchart may be viewed as method steps and / or actions resulting from the operation of computer program code and / or as a plurality of coupled logic circuit elements constructed to perform associated functions (one or more). For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium. The computer program comprises program code configured to perform the methods described above.

[0110] In the context of this disclosure, a machine-readable medium can be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any preferred combination thereof. More specific examples of machine-readable storage media include electrical connections with one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any preferred combination thereof.

[0111] Computer program code for performing the methods of this disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to the processor of a general-purpose computer, a dedicated computer, or other programmable data processing device, and when executed by the computer processor or other programmable data processing device, the program code will perform the functions / operations specified in the flowcharts and / or block diagrams. The program code may run entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.

[0112] Furthermore, although the operations are described in a specific order, this should not be understood as requiring that such operations be performed in a specific illustrated or sequential order, or that all illustrated operations be performed to achieve a desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limitations on the scope of any invention or any claims, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described herein in the context of separate embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately in multiple embodiments or in any preferred subcombination.

[0113] It should be noted that the specification and drawings are merely illustrative of the principles of the proposed method and apparatus. Therefore, those skilled in the art will understand that various configurations embodying the principles of the invention and falling within its spirit and scope, although not expressly described or illustrated herein, may be conceived. Furthermore, all examples described herein are explicitly intended solely for educational purposes to assist the reader in understanding the principles of the proposed method and apparatus, and the concepts to which the inventors contribute to advancing the art, and are to be interpreted without limitation to the examples and conditions thus specifically described. Moreover, all statements herein describing the principles, aspects, and embodiments of the invention, and their individual examples, are intended to encompass their equivalents.

[0114] Several aspects are described below. [Aspect 1] A method for encoding directional audio data: A step of receiving a mono audio signal corresponding to an audio object and a representation of a radiation pattern corresponding to the audio object, wherein the radiation pattern includes a step of sound levels corresponding to a plurality of sample times, a plurality of frequency bands, and a plurality of directions; The step of encoding the aforementioned mono audio signal; The step includes encoding the source radiation pattern to determine the radiation pattern metadata, Encoding the radiation pattern includes determining a spherical harmonic transform of the representation of the radiation pattern, compressing the spherical harmonic transform, and obtaining encoded radiation pattern metadata. method. [Aspect 2] The method according to embodiment 1, further comprising encoding multiple directional audio objects based on a cluster of audio objects, wherein the radiation pattern represents a centroid that reflects the average sound level value for each frequency band. [Aspect 3] The method according to aspect 2, wherein the plurality of directional audio objects are encoded as a single directional audio object having a directivity corresponding to a time-varying, energy-weighted average of the spherical harmonic coefficients of each audio object. [Aspect 4] The method according to embodiment 2 or 3, wherein the encoded radiant pattern metadata indicates the location of a cluster of audio objects, which is the average of the positions of each audio object. [Aspect 5] The method according to any one of embodiments 1 to 4, further comprising encoding group metadata relating to the emission pattern of a group of directional audio objects. [Aspect 6] The method according to any one of embodiments 1 to 5, wherein the source radiation pattern is rescaled with respect to the amplitude of the input radiation pattern in a certain direction for each frequency to determine a normalized radiation pattern. [Aspect 7] The method according to any one of embodiments 1 to 6, wherein compressing the spherical harmonic transform includes at least one of singular value decomposition, principal component analysis, discrete cosine transform, data-independent basis, or eliminating spherical harmonic coefficients of the spherical harmonic transform that are above a threshold order of spherical harmonic coefficients. [Aspect 8] A method for decoding audio data: The process involves receiving the encoded core audio signal, encoded radiation pattern metadata, and encoded audio object metadata; The steps include: decoding the encoded core audio signal to determine the core audio signal; The steps include: decoding the encoded radiation pattern metadata to determine the decoded radiation pattern; The step of decoding the aforementioned audio object metadata; The step includes rendering the core audio signal based on the audio object metadata and the decoded emission pattern, method. [Aspect 9] The method according to embodiment 8, wherein the audio object metadata includes at least one time-varying 3-degree-of-freedom (3DoF) or 6-degree-of-freedom (6DoF) source orientation information. [Aspect 10] The method according to embodiment 8 or 9, wherein the core audio signal includes a plurality of directional objects based on a cluster of objects, and the decoded radiation pattern represents a centroid that reflects the average value for each frequency band. [Aspect 11] The method according to any one of embodiments 8 to 10, wherein the rendering is at least in part based on applying subband gain to the decoded core audio signal based on the decoded radiated data. [Aspect 12] The method according to any one of embodiments 8 to 11, wherein the encoded radiation pattern metadata corresponds to a set of spherical harmonic coefficients that vary with time and frequency. [Aspect 13] The method according to any one of embodiments 8 to 12, wherein the encoded radiation pattern metadata includes audio object type metadata. [Aspect 14] The method according to embodiment 13, wherein the audio object type metadata indicates parametric directional pattern data, and the parametric directional pattern data includes one or more functions selected from a list of functions consisting of cosine functions, sine functions, or cardioid functions. [Aspect 15] The method according to aspect 13, wherein the audio object type metadata indicates database directional pattern data, and decoding the encoded radiant pattern metadata to determine the decoded radiant pattern involves querying a directional data structure containing the audio object type and the corresponding directional pattern data. [Aspect 16] The method according to embodiment 13, wherein the audio object type metadata represents dynamic directional pattern data, and the dynamic directional pattern data corresponds to a set of spherical harmonic coefficients that vary with time and frequency. [Aspect 17] The method according to embodiment 16, further comprising receiving the dynamic directional pattern data before receiving the encoded core audio signal. [Aspect 18] Interface system; and An audio decoding device having a control system, The aforementioned control system: A step of receiving audio data corresponding to at least one audio object via the interface system, wherein the audio data includes a monophonic audio signal, audio object position metadata, audio object size metadata, and rendering parameters; The system is configured to perform the steps of: determining whether the rendering parameter indicates a position mode or a directional mode; and, if it is determined that the rendering parameter indicates a directional mode, rendering the audio data for playback through at least one loudspeaker according to a directional pattern indicated by at least one of the position metadata or the size metadata. Device. [Aspect 19] The apparatus according to embodiment 18, wherein rendering the audio data includes interpreting the audio object position metadata as audio object orientation metadata. [Aspect 20] The apparatus according to embodiment 19, wherein the audio object position metadata includes at least one of x,y,z coordinate data, spherical coordinate data, or cylindrical coordinate data, and the audio object orientation metadata includes yaw, pitch, and roll data. [Aspect 21] The apparatus according to any one of embodiments 18 to 20, wherein rendering of the audio data includes interpreting the audio object size metadata as directional metadata corresponding to the directional pattern. [Aspect 22] The apparatus according to any one of embodiments 18 to 21, wherein rendering the audio data includes querying a data structure containing a plurality of directional patterns and mapping at least one of the position metadata or the size metadata to one or more of the directional patterns. [Aspect 23] The apparatus according to embodiment 22, wherein the control system is configured to receive the data structure via the interface system. [Aspect 24] The apparatus according to embodiment 23, wherein the data structure is received prior to the audio data. [Aspect 25] The apparatus according to any one of embodiments 18 to 24, wherein the aforementioned audio data is received in Dolby Atmos format. [Aspect 26] The device according to any one of embodiments 18 to 25, wherein the audio object position metadata corresponds to world coordinates or model coordinates.

Claims

1. Interface system; and An audio decoding device having a control system, The aforementioned control system is: Steps include: receiving audio data corresponding to at least one audio object via the interface system, wherein the audio data includes a monophonic audio signal, audio object position metadata, and rendering parameters, and the audio object position metadata includes 6DoF source orientation information; The system is configured to perform the steps of: determining whether the rendering parameter indicates a position mode or a directional mode; and, if it is determined that the rendering parameter indicates a directional mode, rendering the audio data for playback through at least one loudspeaker according to a directional pattern indicated by at least one of the audio object position metadata. Device.

2. A method for decoding audio data, performed by an audio decoding device, the method being: Steps include: receiving audio data corresponding to at least one audio object via an interface system, wherein the audio data includes a monophonic audio signal, audio object position metadata, and rendering parameters, and the audio object position metadata includes 6DoF source orientation information; The steps include: determining whether the rendering parameter indicates a position mode or a directional mode; and, if it is determined that the rendering parameter indicates a directional mode, rendering the audio data for playback through at least one loudspeaker according to a directional pattern indicated by at least one of the audio object position metadata, method.

3. A non-temporary computer-readable medium that stores instructions, when executed by one or more processors, causing the one or more processors to perform the method described in Claim 2.

Citation Information

Patent Citations

  • Sound signal processing method and sound field reproduction system

    JP2007043334A

  • Obtaining sparseness information for higher-order ambisonic audio renderers

    JP2017520177A

  • Scalable downmix design with feedback for object-based surround codec

    US20140023196A1

  • Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding

    WO2019068638A1