Method, device and system for encoding and decoding directional sound source
By encoding audio data with spherical harmonic transforms and compression techniques, the patent addresses the challenge of representing complex sound source directivity in interactive environments, achieving efficient and realistic sound rendering in virtual reality applications.
Patent Information
- Application Number
- JP2025064049
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-10-04
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-04-15
AI Technical Summary
Existing audio game engines and virtual reality applications struggle to accurately represent complex radiation patterns of sound sources, particularly in interactive environments, as they rely on simplistic directional gains and high-frequency roll-off filters, failing to capture real-world sound source directivity and higher-order acoustic coupling.
Encoding audio data using spherical harmonic transforms to represent radiation patterns, which are compressed and encoded as metadata, allowing for dynamic rendering of multiple sound sources as clusters or groups, with methods involving singular value decomposition and principal component analysis for efficient data reduction.
Enables realistic and efficient rendering of complex sound scenes with multiple dynamic sound sources, maintaining perceived quality and reducing bitrate requirements for transmission.
Smart Images

Figure 2025106455000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Patent Application No. 62 / 658,067, filed Aug. 16, 2018; U.S. Patent Application No. 62 / 681,429, filed Jun. 6, 2018; and U.S. Patent Application No. 62 / 741,419, filed Oct. 4, 2018. The contents of these applications are hereby incorporated by reference in their entirety.
[0002] Technical Field The present disclosure relates to the encoding and decoding of directional sound sources and auditory scenes based on multiple dynamic and / or moving directional sound sources.
Background Art
[0003] Real - world sound sources, whether natural or artificial (speakers, musical instruments, voices, mechanical devices), radiate sound in a non - isotropic manner. Characterizing the radiation pattern (or “directivity”) of a sound source can be crucial for proper rendering, especially in the context of interactive environments such as video games and virtual reality / augmented reality (VR / AR) applications. In these environments, users generally interact with directional audio objects by walking around them, thereby changing the auditory perspective regarding the generated sound (also called 6 - degree - of - freedom [DoF] rendering). Users can also grab and dynamically rotate virtual objects, which also requires rendering in different directions in the radiation pattern of the corresponding sound source(s). In addition to a more realistic rendering of the direct propagation effect from the source to the listener, the radiation characteristics also play a major role in the higher - order acoustic coupling between the source and its environment (e.g., the virtual environment in a game), and thus affect the reverberant sound (i.e., waves that bounce back and forth as in an echo). As a result, such reverberation can affect other spatial cues such as the perceived distance.
[0004] Most audio game engines provide some way to represent and render directional sound sources, but generally are limited to simple directional gains that rely on the definition of a simple first-order cosine function or "sound cone" (e.g., a power cosine function) and a simple high-frequency roll-off filter. These representations are insufficient to represent real-world radiation patterns and are also not well-suited to simplified / combined representations of multiple directional sound sources. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0005] Various audio processing methods are disclosed herein. Some such methods may involve encoding directional audio data. For example, some methods may involve receiving a mono audio signal corresponding to an audio object and a representation of a radiation pattern corresponding to the audio object. The radiation pattern may include, for example, sound levels corresponding to a plurality of sample times, a plurality of frequency bands, and a plurality of directions. Some such methods may involve encoding the mono audio signal and encoding the source radiation pattern to determine radiation pattern metadata. Encoding the radiation pattern may involve determining a spherical harmonic transform of the representation of the radiation pattern and compressing the spherical harmonic transform to obtain encoded radiation pattern metadata.
[0006] Some such methods may involve encoding multiple directional audio objects based on clusters of audio objects. The radiation pattern may represent a centroid that reflects the average sound level value for each frequency band. In some such implementations, the multiple directional audio objects are encoded as a single directional audio object having a directivity corresponding to a time-varying, energy-weighted average of the spherical harmonic function coefficients of each audio object. The encoded radiation pattern metadata may indicate the location of a cluster of audio objects that is the average of the locations of each audio object.
[0007] Some methods may involve encoding group metadata regarding the radiation pattern of a group of directional audio objects. In some examples, the source radiation pattern may be rescaled relative to the amplitude of the input radiation pattern in a certain direction for each frequency to determine a normalized radiation pattern. According to some implementations, compressing the spherical harmonic function transform may involve singular value decomposition, principal component analysis, discrete cosine transform, data-independent bases, and / or eliminating spherical harmonic function coefficients of the spherical harmonic function transform above a threshold order of the spherical harmonic function coefficients.
[0008] Some alternative methods may be involved in decoding audio data. For example, some such methods may be involved in receiving an encoded core audio signal, encoded radiation pattern metadata, and encoded audio object metadata, and decoding the encoded core audio signal to determine the core audio signal. Some such methods may be involved in decoding the encoded radiation pattern metadata to determine the decoded radiation pattern, decoding the audio object metadata, and rendering the core audio signal based on the audio object metadata and the decoded radiation pattern.
[0009] In some cases, the audio object metadata may include at least one of time-varying 3 degrees of freedom (3DoF) or 6 degrees of freedom (6DoF) source orientation information. The core audio signal may include a plurality of directional objects based on a cluster of objects. The decoded radiation pattern may represent a centroid reflecting an average value for each frequency band. In some examples, the rendering may be based at least in part on applying subband gains based on the decoded radiation data to the decoded core audio signal. The encoded radiation pattern metadata may correspond to a time- and frequency-varying set of spherical harmonic function coefficients.
[0010] According to some implementations, the encoded radiation pattern metadata may include audio object type metadata. The audio object type metadata may indicate, for example, parametric directivity pattern data. The parametric directivity pattern data may include cosine functions, sine functions, and / or cardioid functions. In some examples, the audio object type metadata may indicate database directivity pattern data. Determining the decoded radiation pattern by decoding the encoded radiation pattern metadata may involve querying a directivity data structure that includes audio object type and corresponding directivity pattern data. In some examples, the audio object type metadata may indicate dynamic directivity pattern data. The dynamic directivity pattern data may correspond to a time- and frequency-varying set of spherical harmonic function coefficients. Some methods may involve receiving the dynamic directivity pattern data before receiving the encoded core audio signal.
[0011] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc., of the type described herein. Thus, various innovative aspects of the subject matter disclosed herein can be implemented on one or more non-transitory media storing software. The software may include instructions for controlling at least one device to process audio data, for example. The software may be executable by one or more components of a control system as disclosed herein, for example. The software may include instructions for performing one or more of the methods disclosed herein, for example.
[0012] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more apparatuses may be configured to at least partially execute the methods disclosed herein. In some implementations, the apparatus may include an interface system and a control system. The interface system may include one or more network interfaces, one or more interfaces between the control system and a memory system, one or more interfaces between the control system and another device, and / or one or more external device interfaces. The control system may include at least one of a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gates or transistor logic, or discrete hardware components. Thus, in some implementations, the control system may include one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors.
[0013] According to some such examples, the control system may be configured to receive audio data corresponding to at least one audio object via the interface system. In some examples, the audio data may include a monophonic audio signal, audio object position metadata, audio object size metadata, and rendering parameters. Some such methods may determine whether the rendering parameters indicate a positional mode or a directional mode, and, upon determining that the rendering parameters indicate a directional mode, may involve rendering the audio data for playback via at least one loudspeaker according to a directional pattern indicated by the position metadata and / or the size metadata.
[0014] In some examples, the rendering of audio data may involve interpreting audio object position metadata as audio object orientation metadata. The audio object position metadata may include, for example, x, y, z coordinate data, spherical coordinate data, and / or cylindrical coordinate data. In some cases, the audio object orientation metadata may include yaw, pitch, roll data.
[0015] According to some examples, the rendering of audio data may involve interpreting audio object size metadata as directivity metadata corresponding to a directivity pattern. In some implementations, the rendering of audio data may involve querying a data structure that includes multiple directivity patterns and mapping position metadata and / or size metadata to one or more of the directivity patterns. In some cases, the control system may be configured to receive the data structure via an interface system. In some examples, the data structure may be received prior to the audio data. In some implementations, the audio data may be received in a Dolby Atmos format. The audio object position metadata may correspond to, for example, world coordinates or model coordinates.
[0016] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, the drawings, and the claims. Note that the relative dimensions of the figures below may not be drawn to scale. Like reference numerals and symbols in the various drawings generally denote like elements.
Brief Description of the Drawings
[0017]
Figure 1A
[0018]
Figure 1B
[0019]
Figure 1C
[0020]
Figure 2
[0021]
Figure 2C
[0022]
Figure 3
[0023]
Figure 4
[0024]
Figure 5A
[0025]
Figure 5B
[0026]
Figure 6
[0027]
Figure 7
[0028]
Figure 8
[0029] Like reference numerals and characters in the various drawings indicate like elements. DETAILED DESCRIPTION
[0030] One aspect of the present disclosure relates to the representation of complex radiation patterns and efficient encoding. Some such implementations may include one or more of the following: 1. Representation of a general sound radiation pattern as time- and frequency-dependent Nth-order coefficients of a real-valued spherical harmonics (SPH) decomposition (N≥1). This representation can also be extended depending on the level of the reproduced audio signal. Contrary to the case where the directional source signal itself is in a PCM representation such as HOA, the mono-object signal can be encoded separately from its directivity information, and the directivity information is represented as a set of time-dependent scalar SPH coefficients in the respective subbands. 2. An efficient encoding scheme to reduce the bitrate required to represent this information 3. A solution for dynamically combining radiation patterns such that a scene composed of several radiating sound sources can be represented by an equivalent reduced number of sources while maintaining its perceived quality during rendering.
[0031] One aspect of the present disclosure relates to representing a general radiation pattern to complement the metadata for each mono-audio object with a set of time / frequency-dependent coefficients representing the directivity of the mono-audio object projected in a spherical harmonics basis of order N (N≥1).
[0032] The primary radiation pattern can be represented by a set of four scalar gain coefficients for a predefined set of frequency bands (e.g., 1 / 3 octave). The set of frequency bands is also known as bins or sub-bands. The bins or sub-bands may be determined based on a perceptual filter bank for a short-time Fourier transform (STFT) or a single data frame (e.g., 512 samples as in Dolby Atmos). The resulting pattern can be rendered by evaluating a spherical harmonic decomposition in the required directions around the object.
[0033] Generally, this radiation pattern is a characteristic of the source and may be constant over time. However, it may be beneficial to update this set of coefficients at regular time intervals to represent a dynamic scene where the object rotates or changes, or to ensure that the data can be accessed randomly. In the context of a dynamic auditory scene with a moving object, the result of object rotation can be directly encoded in the time-varying coefficients without the need for an explicit separate encoding of the object orientation.
[0034] Each type of sound source typically has a characteristic radiation / emission pattern that varies by frequency band. For example, a violin may have a very different radiation pattern from a trumpet, drum, or bell. Additionally, a sound source such as an instrument may radiate differently at pianissimo and fortissimo performance levels. As a result, the radiation pattern may be a function not only of the direction around the sound-emitting object but also of the pressure level of the radiated audio signal, and the pressure level may also vary over time.
[0035] Thus, instead of simply representing a sound field at a point in space, some implementations involve encoding audio data corresponding to the radiation pattern of an audio object so that it can be rendered from different vantage points. In some cases, the radiation pattern may be a radiation pattern that varies with time and frequency. The audio data input into the encoding process may, in some cases, include multiple channels of audio data from a directional microphone (e.g., 4, 6, 8, 20 or more channels). Each channel may correspond to data from a microphone at a specific location in the space around the sound source, from which the radiation pattern can then be derived. Given that the relative direction from each microphone to the sound source is known, this can be achieved by a numerical fitting of a set of spherical harmonic function coefficients such that the resulting spherical function best matches the observed energy levels in the various subbands of each input microphone signal. See, for example, the methods and systems described in International Application No. PCT / US2017 / 053946, “Method, Systems and Apparatus for Determining Audio Representations” by Nicolas Tsingos and Pradeep Kumar Govindaraju. This application is hereby incorporated by reference. In other examples, the radiation pattern of the audio object may be determined by numerical simulation.
[0036] Instead of simply encoding audio data from a directional microphone at the sample level, some implementations involve encoding a monophonic audio object signal along with corresponding radiation pattern metadata that represents the radiation pattern for at least some of the encoded audio objects. In some implementations, the radiation pattern metadata can be expressed as spherical harmonic function data. Some such implementations may involve a smoothing process and / or a compression / data reduction process.
[0037] Figure 1A is a flowchart showing blocks of an audio encoding method according to an example. Method 1 may be implemented, for example, by a control system (such as control system 815 described below with reference to FIG. 8) including one or more processors and one or more non-transitory memory devices. As with the other disclosed methods, not all blocks of Method 1 are necessarily executed in the order shown in FIG. 1A. Further, alternative methods may include more or fewer blocks.
[0038] In this example, block 5 involves receiving a mono audio signal corresponding to an audio object and also receiving a representation of the radiation pattern corresponding to the audio object. According to this implementation, the radiation pattern includes sound levels corresponding to multiple sample times, multiple frequency bands, and multiple directions. According to this example, block 10 involves encoding the mono audio signal.
[0039] In the example shown in FIG. 1A, block 15 is involved in encoding a source radiation pattern to determine radiation pattern metadata. According to this implementation, encoding the representation of the radiation pattern involves determining the spherical harmonic function transform of the representation of the radiation pattern and compressing the spherical harmonic function transform to obtain the encoded radiation pattern metadata. In some implementations, the representation of the radiation pattern may be rescaled for each frequency with respect to the amplitude of the input radiation pattern in a certain direction to determine a normalized radiation pattern.
[0040] In some cases, compressing the spherical harmonic function transform may involve discarding some higher-order spherical harmonic function coefficients. Some such examples involve removing the spherical harmonic function coefficients of the spherical harmonic function transform above a threshold order of the spherical harmonic function coefficients, such as above order 3, above order 4, above order 5.
[0041] However, some implementations may involve alternative and / or additional compression methods. According to some such implementations, compressing the spherical harmonic function transform may involve singular value decomposition, principal component analysis, discrete cosine transform, data-independent bases, and / or other methods.
[0042] According to some examples, method 1 may also involve encoding a plurality of directional audio objects as a group or “cluster” of audio objects. Some implementations may involve encoding group metadata regarding the radiation pattern of a group of directional audio objects. In some cases, a plurality of directional audio objects may be encoded as a single directional audio object having a directivity corresponding to an energy-weighted average of the time-varying spherical harmonic function coefficients of each audio object. In some such examples, the encoded radiation pattern metadata may represent a centroid corresponding to the average sound level value for each frequency band. For example, the encoded radiation pattern metadata (or related metadata) may indicate the position of a cluster of audio objects that is the average of the positions of each directional audio object within the cluster.
[0043] FIG. 1B shows blocks of a process that may be implemented by encoding system 100 to dynamically encode frame-by-frame directivity information for a directional audio object, according to one example. This process may be implemented via a control system, such as control system 815 described below with reference to FIG. 8, for example. Encoding system 100 may receive a mono audio signal 101 that may correspond to a mono object signal as discussed above. Mono audio signal 101 may be encoded at block 111 and provided to serialization block 112.
[0044] At block 102, static or time-varying directional energy samples at different sound levels in a set of frequency bands with respect to a reference coordinate system may be processed. The reference coordinate system may be determined in a certain coordinate space, such as a model coordinate space or a world coordinate space.
[0045] In block 105, frequency-dependent rescaling of the time-varying directional energy samples from block 102 may be performed. In one example, the frequency-dependent rescaling may be performed according to the example shown in FIGS. 2A - 2C as described below. The normalization may be based on, for example, the rescaling of the amplitude in the low-frequency direction for high frequencies.
[0046] The frequency-dependent rescaling may be renormalized based on the assumed capture direction of the core audio. Such an assumed capture direction of the core audio may represent the listening direction with respect to the sound source. For example, this listening direction may be called the gaze direction, where the gaze direction may be in a certain direction (e.g., the forward direction or the backward direction) with respect to the coordinate system.
[0047] In block 106, the rescaled directivity output of 105 may be projected onto a spherical harmonic function basis, and as a result, the coefficients of the spherical harmonic functions are given.
[0048] In block 108, the spherical coefficients of block 106 are processed based on the instantaneous sound level 107 and / or information from the rotation block 109. The instantaneous sound level 107 may be measured at a certain time in a certain direction. The information from the rotation block 109 may indicate the (optional) rotation of the time-varying source orientation 103. In one example, in block 109, the spherical coefficients can be adjusted to account for the time-dependent modification of the source orientation with respect to the originally recorded input data.
[0049] In block 108, target level determination may be further performed based on equalization determined with respect to the direction of the assumed capture direction of the core audio signal. Block 108 may output a set of rotated spherical coefficients equalized based on the target level determination.
[0050] In block 110, the encoding of the radiation pattern may be based on the projection of the spherical coefficients associated with the source radiation pattern onto a smaller subspace, and as a result, encoded radiation pattern metadata is obtained. As shown in FIG. 1A, in block 110, SVD decomposition and a compression algorithm may be performed on the spherical coefficients output by block 108. In one example, the SVD decomposition and compression algorithm of block 110 may be performed according to the principles described in connection with equations 11-13 below.
[0051] Alternatively, block 110 may involve using principal component analysis (PCA) and / or other methods such as data-independent bases, such as 2D DCT, to project the spherical harmonic function representation into a space that leads to irreversible compression. The output of 110 may be a matrix T representing the projection of the input data onto a smaller subspace, i.e., the encoded radiation pattern T. The encoded radiation pattern T, the encoded core mono-audio signal 111, and any other object metadata 104 (e.g., x, y, z, arbitrary source orientation, etc.) may be serialized in serialization block 112 to output an encoded bitstream. In some examples, the radiation structure may be represented by the following bitstream syntax structure in each encoded audio frame:
Number
[0052] Such syntax may include different sets of coefficients for different pressure / intensity levels of the sound source. Alternatively, if the directivity information is available at different signal levels and the level of the source cannot be further determined during playback, a single set of coefficients may be generated dynamically. For example, such coefficients may be generated by interpolating between low-level coefficients and high-level coefficients based on the time-varying level of the object audio signal during encoding.
[0053] The input radiation pattern for a mono audio object signal may be "normalized" with respect to a given direction such as the main response axis (which may be the recorded direction or the average of multiple recordings), and the encoded directivity and the final rendering may need to be consistent with this "normalization". In one example, this normalization may be specified as metadata. Generally, it is desirable to encode the core audio signal that would convey a good representation of the timbre of the object in the absence of applied directivity information.
[0054] Directional encoding One aspect of the present disclosure is directed to implementing an efficient encoding scheme for directivity information since the number of coefficients increases quadratically with the degree of decomposition. An efficient encoding scheme for directivity information may be implemented, for example, for the final transmission of the auditory scene to an end rendering device over a network with limited bandwidth.
[0055] Assuming 16 bits are used to represent each coefficient, a fourth-order spherical harmonic function representation in a 1 / 3 octave band would require approximately 25×31 = 12 kbit per frame. To refresh this information at 30 Hz, a transmission bitrate of at least 400 kbps is required, which is more than what current object-based audio codecs currently require to transmit both audio and object metadata. In one example, the radiation pattern is G(θ i ,φi , ω) Equation (1) may be represented by
[0056] In Equation (1), (θ i , φ i ), i ∈ {1…P} represent the discrete co-latitude angle θ ∈ [0, π] and azimuth angle φ ∈ [0, 2π) with respect to the sound source, P represents the total number of discrete angles, and ω represents the spectral frequency. FIGS. 2A and 2B represent the radiation patterns of an audio object in two different frequency bands. FIG. 2A may represent, for example, the radiation pattern of an audio object in a frequency band of 100 to 300 Hz, and FIG. 2B may represent, for example, the radiation pattern of the same audio object in a frequency band of 1 kHz to 2 kHz. Since low frequencies tend to be relatively more omnidirectional, the radiation pattern shown in FIG. 2A is relatively closer to circular than the radiation pattern shown in FIG. 2B. In FIG. 2A, G(θ0, φ0, ω) represents the radiation pattern in the direction of the main response axis 200, while G(θ1, φ1, ω) represents the radiation pattern in an arbitrary direction 205.
[0057] In some examples, the radiation pattern may be captured and determined by a plurality of microphones physically arranged around the sound source corresponding to the audio object, but in other examples, the radiation pattern may be determined via numerical simulation. In the example of a plurality of microphones, the radiation pattern may change over time, for example, reflecting a live recording. The radiation pattern may be captured at various frequencies including low frequencies (e.g., <100 Hz), mid frequencies (100 Hz < and > 1 kHz), and high frequencies (> 10 kHz). The radiation pattern is sometimes also known as a spatial representation.
[0058] In another example, the radiation pattern may reflect normalization based on the captured radiation pattern G(θ i , φ i , ω) at a certain frequency in a certain direction. For example:
Number
[0059] In Equation (2), G(θ0, φ0, ω) represents the radiation pattern in the direction of the main response axis. Referring again to FIG. 2B, in one example, the radiation pattern G(θ i , φ i , ω) and the normalized radiation pattern H(θ i , φ i , ω) can be seen. FIG. 2C is a graph showing an example of the normalized radiation pattern and the non-normalized radiation pattern according to one example. In this example, the normalized radiation pattern in the direction of the main response axis, represented as H(θ0, φ0, ω) in FIG. 2C, has substantially the same amplitude over the illustrated range of the frequency band. In this example, the normalized radiation pattern in direction 205 (shown in FIG. 2A), represented as H(θ1, φ1, ω) in FIG. 2C, has a relatively higher amplitude at higher frequencies than the non-normalized radiation pattern represented as G(θ1, φ1, ω) in FIG. 2C. For a given frequency band, the radiation pattern may be assumed to be constant for notational convenience, but in practice, it may change over time, for example, by different bowing techniques used in stringed instruments.
[0060] The radiation pattern, or its parametric representation, may be transmitted. Preprocessing of the radiation pattern may be performed prior to its transmission. In one example, the radiation pattern or parametric representation may be preprocessed by a computational algorithm, an example of which is shown in connection with FIG. 1A. After preprocessing, the radiation pattern may be decomposed, for example,
Number
[0061] In Equation (3), H(θ i , φ i , ω) represents the spatial representation,
Number
Mathematics
Mathematics
[0062] In Equation (4), P n m (x) is the associated Legendre polynomial, the order m ∈ {-N…N}, and the degree n ∈ {0…N},
Mathematics
[0063] Other spherical bases may be used. Any method for performing spherical harmonic function conversion on discrete data may be used. In one example, the matrix conversion
Mathematics
Mathematics
[0064] In Equation (7),
Mathematics
[0065] Pseudo-inverse matrix
Mathematics
[0066] The regularized solution may be applicable even when the distribution of spherical samples contains a large amount of missing data. The missing data may correspond to regions or directions where directional samples are not available (e.g., because the coverage of the microphone is non-uniform). In many cases, the distribution of spatial samples is sufficiently uniform that an identity weighting matrix W yields acceptable results. Also, often P ≫ (N + 1) 2 is assumed, and the spherical harmonic function representation [Number] is the spatial representation [Number] contains fewer elements than
[0067] and thereby provides a first stage of irreversible compression that smooths the radiation pattern data. k Here, consider discrete frequency bands ω [Number]
[0068] That is, the spatial representation H(ω) can be determined based on frequency bins / bands / sets. As a result, the spherical harmonic function representation may be: [Number] based on
[0069] In Equation (10),
Number
Number
Number
Number
[0070] Some embodiments may involve performing a singular value decomposition (SVD), where
Number
Number
Number
[0071] Let O=(N + 1) 2 In some examples, to achieve compression, the encoder
Number
[0072] In Equation (12),
Number
Number
Number
Number
[0073] The following are three examples for transmitting the truncated decomposition vectors and the truncated right singular vectors: 1. The transmitter may transmit the encoded radiation T and the truncated right singular vector V' independently for each object. 2. The objects may be grouped, for example, according to a similarity metric, and U and V may be calculated as representative bases for multiple objects. Thus, the encoded radiation can be transmitted for each object, and U and V can be transmitted for each group of objects. 3. The left and right singular matrices U and V may be calculated in advance on a large database of representative data (e.g., training data), and information about V may be stored at the receiver side. In some such examples, only the encoded radiation may be transmitted for each object. The DCT is another example of a basis that may be stored at the receiver side.
[0074] Spatial encoding of directional objects When a complex auditory scene containing multiple objects is encoded and transmitted, a spatial encoding technique in which individual objects are replaced with a smaller number of representative clusters can be applied in a way that best preserves the auditory perception of the scene. Generally, replacing a group of sound sources with a representative "centroid" requires calculating an aggregated / mean value for each metadata field. For example, the position of a cluster of sound sources can be the average of the positions of each sound source. (See, for example, Equations 1-12) By expressing the radiation pattern of each source using spherical harmonic function decomposition as described above, it is possible to linearly combine the set of coefficients in each subband for each source to construct an average radiation pattern for the cluster of sources. By calculating an average weighted over time by loudness or the energy of the spherical harmonic function coefficients, it is possible to construct a perceptually optimized representation that better preserves the original scene over time.
[0075] Figure 1C shows blocks of a process that can be implemented by a decoding system according to an example. The blocks shown in Figure 1C may be implemented, for example, by a control system of a decoding device (such as control system 815 described below with reference to Figure 8) that includes one or more processors and one or more non-transitory memory devices. In block 150, metadata and an encoded core mono audio signal may be received and deserialized. The deserialized information may include object metadata 151, the encoded core audio signal, and the encoded spherical coefficients. In block 152, the encoded core audio signal may be decoded. In block 153, the encoded spherical coefficients may be decoded. The encoded radiation pattern information may include the encoded radiation pattern T and / or matrix V. Matrix V is,
Number
[0076] The object metadata 151 may include information regarding the relative direction from the source to the listener. In one example, the metadata 151 may include information regarding the distance and direction of the listener and the distance and direction of one or more objects relative to a 6DoF space. For example, the metadata 151 may include information regarding the relative rotation, distance, and direction of the source in a 6DoF space. In the example of multiple objects within a cluster, the metadata field may reflect information regarding a representative "centroid" that reflects the aggregated value / average value of the cluster of objects.
[0077] Next, the renderer 154 may render the decoded core audio signal and the decoded spherical harmonic function coefficients. In one example, the renderer 154 may render the decoded core audio signal and the decoded spherical harmonic function coefficients based on the object metadata 151. The renderer 154 may determine a subband gain for the spherical coefficients of the radiation pattern based on information from the metadata 151, for example, the relative direction from the source to the listener. The renderer 154 may then render the core audio object signal based on the determined subband gain for the corresponding decoded radiation pattern(s), the source and / or listener pose information (e.g., x, y, z, yaw, pitch, roll) 155. The listener pose information may correspond to the position and viewing direction of the user in a 6DoF space. The listener pose information may be received from a source local to the VR playback system, such as an optical tracking device. The source pose information corresponds to the position and orientation in space of the object emitting the sound. It can also be inferred from a local tracking system. For example, when the user's hand is tracked and interactively manipulates a virtual sound-emitting object, or when a tracked physical prop / proxy object is used.
[0078] Figure 3 shows an example of a hierarchy that includes audio data and various types of metadata. Similar to the other drawings provided herein, the number and types of audio data and metadata shown in Figure 3 are provided merely as examples. Some encoders may provide the complete set of audio data and metadata shown in Figure 3 (data set 345), while other encoders may provide only a portion of the metadata shown in Figure 3, for example, only data set 315, only data set 325, or only data set 335.
[0079] In this example, the audio data includes a monophonic audio signal 301. The monophonic audio signal 301 is an example of what may sometimes be referred to herein as the "core audio signal", although in some examples, the core audio signal may include various audio signals corresponding to a plurality of audio objects included in a cluster.
[0080] In this example, the audio object position metadata 305 is expressed as Cartesian coordinates. However, in alternative examples, the audio object position metadata 305 may be expressed via other types of coordinates such as spherical coordinates or polar coordinates. Thus, the audio object position metadata 305 may include three degrees of freedom (3DoF) position information. According to this example, the audio object metadata includes audio object size metadata 310. In alternative examples, the audio object metadata may include one or more other types of audio object metadata.
[0081] In this implementation, the dataset 315 includes the monophonic audio signal 301, the audio object position metadata 305, and the audio object size metadata 310. The dataset 315 may be provided, for example, in the Dolby Atmos (trademark) audio data format.
[0082] In this example, the dataset 315 also includes optional rendering parameter R. According to some disclosed implementations, the optional rendering parameter R can indicate whether at least a part of the audio object metadata of the dataset 315 should be interpreted in its "normal" sense (e.g., as position or size metadata) or as directivity metadata. In some disclosed implementations, the "normal" mode may be referred to herein as the "position mode", and the alternative mode may be referred to herein as the "directivity mode". Some examples are described below with reference to FIGS. 5A - 6.
[0083] According to this example, the orientation metadata 320 includes angular information for representing the yaw, pitch, and roll of the audio object. In this example, the orientation metadata 320 represents yaw, pitch, and roll as Φ, Θ, and Ψ. The dataset 325 includes sufficient information to orient the audio object for a six - degree - of - freedom (6 DoF) application.
[0084] In this example, the dataset 335 includes audio object type metadata 330. In some implementations, the audio object type metadata 330 may be used to indicate the corresponding radiation pattern metadata. The encoded radiation pattern metadata may be used to determine the decoded radiation pattern (e.g., by a decoder or a device that receives audio data from the decoder). In some examples, the audio object type metadata 330 may essentially indicate "I am a trumpet", "I am a violin", etc. In some examples, the decoding device may have access to a database of audio object types and corresponding directivity patterns. According to some examples, the database may be provided together with the encoded audio data or before the transmission of the audio data. Such audio object type metadata 330 may be referred to herein as "data - directivity pattern data".
[0085] According to some examples, the audio object type metadata may indicate parametric directivity pattern data. In some examples, the audio object type metadata 330 may indicate a directivity pattern corresponding to a cosine function of a specified power, or may indicate a cardioid function or the like.
[0086] In some examples, the audio object type metadata 330 may indicate that the radiation pattern corresponds to a set of spherical harmonic function coefficients. For example, the audio object type metadata 330 may indicate that the spherical harmonic function coefficients 340 are provided in the data set 345. In some such examples, the spherical harmonic function coefficients 340 may be a set that varies with time and / or frequency of the spherical harmonic function coefficients, for example, as described above. Such information may require the largest amount of data compared to the rest of the metadata hierarchy shown in FIG. 3. Thus, in some such examples, the spherical harmonic function coefficients 340 may be provided separately from the monophonic audio signal 301 and the corresponding audio object metadata. For example, the spherical harmonic function coefficients 340 may be provided at the start of transmission of the audio data before real-time operation (for example, real-time rendering operations such as games, movies, music performances, etc.) starts.
[0087] According to some implementations, a decoder-side device, such as a device that provides audio to a playback system, may determine the capabilities of the playback system and provide directivity information according to those capabilities. For example, even when the entire data set 345 is provided to the decoder, only the usable portion of the directivity information may be provided to the playback system in some such implementations. In some examples, the decoding device may determine which type(s) of directivity information to use according to the capabilities of the decoding device.
[0088] FIG. 4 is a flow diagram illustrating blocks of an example audio decoding method. Method 400 may be implemented, for example, by a control system (such as control system 815 described below with reference to FIG. 8) of a decoding device that includes one or more processors and one or more non-transitory memory devices. As with the other disclosed methods, not all blocks of method 400 are necessarily executed in the order shown in FIG. 4. Further, alternative methods may include more or fewer blocks.
[0089] In this example, block 405 is involved in receiving an encoded core audio signal, encoded radiation pattern metadata, and encoded audio object metadata. The encoded radiation pattern metadata may include audio object type metadata. The encoded core audio signal may include, for example, a monophonic audio signal. In some examples, the audio object metadata may include 3DoF position information, 6DoF position information and source orientation information, audio object size metadata, etc. The audio object metadata may, in some cases, change over time.
[0090] In this example, block 410 includes decoding the encoded core audio signal to determine a core audio signal. Here, block 415 includes decoding the encoded radiation pattern metadata to determine a decoded radiation pattern. In this example, block 420 is involved in decoding at least a portion of other encoded audio object metadata. Here, block 430 is involved in rendering the core audio signal based on the audio object metadata (such as audio object position, orientation and / or size metadata) and the decoded radiation pattern.
[0091] Block 415 may be involved in various types of operations depending on the specific implementation. In some cases, the audio object type metadata may indicate database-oriented pattern data. Decoding the encoded radiation pattern metadata to determine the decoded radiation pattern may involve querying a directional data structure that includes the audio object type and the corresponding directional pattern data. In some examples, the metadata of the audio object type may indicate parametric directional pattern data, such as directional pattern data corresponding to a cosine function, a sine function, or a cardioid function.
[0092] According to some implementations, the audio object type metadata may indicate dynamic directional pattern data, such as a set that varies over time and / or frequency of spherical harmonic function coefficients. Some such implementations may be involved in receiving the dynamic directional pattern data before receiving the encoded core audio signal.
[0093] In some cases, the core audio signal received at block 405 may include audio signals corresponding to a plurality of audio objects included in the cluster. According to some such examples, the core audio signal may be based on a cluster of audio objects that may include a plurality of directional audio objects. The decoded radiation pattern determined at block 415 may correspond to the centroid of the cluster and may represent an average value for each frequency band of each of the plurality of directional audio objects of the plurality of directional audio objects. The rendering process of block 430 may involve applying a subband gain that is at least partially based on the decoded radiation data to the decoded core audio signal. In some examples, after decoding the core audio signal and decoding and applying a directivity process thereto, the signal may further be virtualized to its intended position with respect to the listener position. For that, known rendering processes such as audio object position metadata and binaural rendering through headphones, rendering using the loudspeakers of the playback environment, etc. are used.
[0094] As described above with reference to FIG. 3, in some implementations, audio data may be accompanied by rendering parameters (shown as R in FIG. 3). The rendering parameters may indicate whether at least some audio object metadata, such as Dolby Atmos metadata, should be interpreted in a normal way (e.g., as position or size metadata) or as directional metadata. The normal mode may sometimes be referred to as the "position mode," and the alternative mode may sometimes be referred to herein as the "directional mode." Thus, in some examples, the rendering parameters may indicate whether to interpret at least some audio object metadata as a direction relative to the speaker or as a position relative to the room or other playback environment. Such implementations may be particularly useful, for example, for directional rendering using smart speakers with multiple drivers, as described below.
[0095] FIG. 5A depicts a drum cymbal. In this example, drum cymbal 505 is shown emitting a sound having a directional pattern 510 with a substantially vertical main response axis 515. The directional pattern 510 itself is also primarily vertical, with some spread from the main response axis 515.
[0096] FIG. 5B shows an example of a speaker system. In this example, speaker system 525 includes a plurality of speakers / transducers configured to radiate sound in various directions, including upward. The top speaker can, in some cases, be used in a conventional Dolby Atmos manner ("position mode") to render the position such that sound is reflected from the ceiling, for example, to simulate a height / ceiling speaker (z = 1). In some such cases, the corresponding Dolby Atmos rendering may include additional height virtualization processing that improves the perception of an audio object having a particular position.
[0097] In other use cases, the same upward-firing speaker(s) can be operated in a "directional mode". This is to simulate, for example, the directional pattern of a drum, cymbal, or other audio object having a directional pattern similar to the directional pattern 510 shown in FIG. 5A. Some speaker systems 525 may be capable of beamforming that can assist in constructing the desired directional pattern. In some examples, no virtualization process is included to reduce the perception of an audio object having a particular position.
[0098] FIG. 6 is a flowchart showing blocks of an example audio method. Method 600 may be implemented, for example, by a control system (such as control system 815 described below with reference to FIG. 8) of a decoding device including one or more processors and one or more non-transitory memory devices. As with other disclosed methods, not all blocks of method 600 are necessarily executed in the order shown in FIG. 6. Further, alternative methods may include more or fewer blocks.
[0099] In this example, block 605 is involved in receiving audio data corresponding to at least one audio object. The audio data includes a monophonic audio signal, audio object position metadata, audio object size metadata, and rendering parameters. In this implementation, block 605 is involved in receiving this data via an interface system (such as interface system 810 of FIG. 8) of the decoding device. In some cases, the audio data may be received in a Dolby Atmos (trademark) format. The audio object position metadata may correspond to world coordinates or model coordinates depending on a particular implementation.
[0100] In this example, block 610 is involved in determining whether the rendering parameter indicates a position mode or a directivity mode. In the example shown in FIG. 6, if it is determined that the rendering parameter indicates the directivity mode, in block 615, the audio data is rendered for playback (for example, playback via at least one loudspeaker, headphones, etc.) according to the directivity pattern indicated by at least one of the position metadata or the size metadata. For example, the directivity pattern may be the same as that shown in FIG. 5A.
[0101] In some examples, the rendering of the audio data may involve interpreting the audio object position metadata as audio object orientation metadata. The audio object position metadata may be Cartesian / x, y, z coordinate data, spherical coordinate data, or cylindrical coordinate data. The audio object orientation metadata may be yaw, pitch, roll metadata.
[0102] According to some implementations, the rendering of the audio data may involve interpreting the audio object size metadata as directivity metadata corresponding to the directivity pattern. In some such examples, the rendering of the audio data may involve querying a data structure that includes a plurality of directivity patterns and mapping at least one of the position metadata or the size metadata to one or more of the directivity patterns. Some such implementations may involve receiving the data structure via an interface system. According to some such implementations, the data structure may be received prior to the audio data.
[0103] FIG. 7 shows an example of encoding a plurality of audio objects. In one example, information 701, 702, 703, etc. of objects 1 to n may be encoded. In one example, in block 710, representative clusters for audio objects 701 to 703 may be determined. In one example, a group of sound sources may be aggregated and represented by a representative "centroid". This involves calculating an aggregated / average value for the metadata field. For example, the position of the sound source cluster can be the average of the positions of each sound source. In block 720, the radiation pattern for the representative cluster can be encoded. In some examples, the radiation pattern for the cluster may be encoded according to the principles described above with reference to FIGS. 1A or 1B.
[0104] FIG. 8 is a block diagram showing an example of components of an apparatus that may be configured to execute at least a portion of the methods disclosed herein. For example, apparatus 805 may be configured to execute one or more of the methods described above with reference to FIGS. 1A - 1C, 4, 6, and / or 7. In some examples, apparatus 805 may be or include a personal computer, a desktop computer, or other local device configured to provide audio processing. In some examples, apparatus 805 may be or include a server. According to some examples, apparatus 805 may be a client device configured to communicate with a server via a network interface. The components of apparatus 805 may be implemented via hardware, via software stored on a non - transitory medium, via firmware, and / or combinations thereof. The types and numbers of components shown in FIG. 8 and other figures disclosed in this application are shown merely by way of example. Alternative implementations may include more, fewer, and / or different components.
[0105] In this example, the apparatus 805 includes an interface system 810 and a control system 815. The interface system 810 may include one or more network interfaces, one or more interfaces between the control system 815 and the memory system, and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). In some implementations, the interface system 810 may include a user interface system. The user interface system may be configured to receive input from a user. In some implementations, the user interface system may be configured to provide feedback to the user. For example, the user interface system may include one or more displays having a corresponding touch and / or gesture detection system. In some examples, the user interface system may include one or more microphones and / or speakers. According to some examples, the user interface system may include a device for providing tactile feedback such as a motor, a vibrator, etc. The control system 815 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gates or transistor logic, and / or discrete hardware components.
[0106] In some examples, the apparatus 805 may be implemented as a single device. However, in some implementations, the apparatus 805 may be implemented as multiple devices. In some such implementations, the functions of the control system 815 may be included in multiple devices. In some examples, the apparatus 805 may be a component of another device.
[0107] Various exemplary embodiments of the present disclosure may be implemented in hardware or special purpose circuitry, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, a microprocessor, or other computing device. In general, it is understood that the present disclosure also encompasses an apparatus suitable for performing the methods described above. For example, an apparatus (a spatial renderer) having a memory and a processor coupled to the memory, where the processor is configured to execute instructions and perform a method according to an embodiment of the present disclosure.
[0108] Various aspects of the exemplary embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation. However, the blocks, apparatus, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, special purpose circuitry or logic, general purpose hardware or a controller, or other computing device, or any combination thereof.
[0109] Furthermore, the various blocks shown in the flowchart may be regarded as method steps and / or acts resulting from operations of computer program code and / or as a plurality of coupled logic circuit elements constructed to perform the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program embodied tangibly on a machine-readable medium. The computer program includes program code configured to perform the above-described method.
[0110] In the context of the present disclosure, a machine-readable medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronics, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium can include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0111] The computer program code for carrying out the methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, and the program code, when executed by the processor of the computer or other programmable data processing apparatus, causes the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer, or entirely on the remote computer or server.
[0112] Furthermore, although the operations are depicted in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve a desired result. In certain situations, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately in multiple embodiments or in any suitable sub-combination.
[0113] It should be noted that the specification and drawings merely illustrate the principles of the proposed methods and apparatuses. Thus, it will be understood that those skilled in the art may devise various configurations that, although not explicitly described or illustrated herein, embody the principles of the invention and are within its spirit and scope. Furthermore, all of the examples described herein are clearly intended solely for the educational purpose of assisting the reader in understanding the principles of the proposed methods and apparatuses and the concepts contributed by the inventors to facilitate the art, and are to be construed without being limited to such specifically described examples and conditions. Moreover, all statements in this specification that describe the principles, aspects, and embodiments of the invention, as well as the individual examples thereof, are intended to encompass their equivalents.
[0114] Some aspects are described below. 〔Aspect 1〕 A method for encoding directional audio data, comprising: Receiving a mono-audio signal corresponding to an audio object and a representation of a radiation pattern corresponding to the audio object, the radiation pattern including sound levels corresponding to a plurality of sample times, a plurality of frequency bands, and a plurality of directions; Encoding the mono-audio signal; Encoding the source radiation pattern to determine radiation pattern metadata, wherein encoding the radiation pattern includes determining a spherical harmonic transform of the representation of the radiation pattern and compressing the spherical harmonic transform to obtain encoded radiation pattern metadata; A method. [Aspect 2] The method according to aspect 1, further comprising encoding a plurality of directional audio objects based on a cluster of audio objects, the radiation pattern representing a centroid reflecting an average sound level value for each frequency band. [Aspect 3] The method according to aspect 2, wherein the plurality of directional audio objects are encoded as a single directional audio object having a directivity corresponding to a time-varying, energy-weighted average of the spherical harmonic function coefficients of each audio object. [Aspect 4] The method according to aspect 2 or 3, wherein the encoded radiation pattern metadata indicates a position of a cluster of audio objects that is an average of the positions of each audio object. [Aspect 5] The method according to any one of aspects 1 to 4, further comprising encoding group metadata regarding a radiation pattern of a group of directional audio objects. [Aspect 6] The method according to any one of aspects 1 to 5, wherein the source radiation pattern is rescaled with respect to an amplitude of an input radiation pattern in a certain direction for each frequency to determine a normalized radiation pattern. [Aspect 7] Compressing the spherical harmonic transform includes at least one of singular value decomposition, principal component analysis, discrete cosine transform, data-independent basis, or eliminating spherical harmonic coefficients of the spherical harmonic transform above a threshold order of the spherical harmonic coefficients, according to any one of aspects 1 to 6 of the method. 〔Aspect 8〕 A method for decoding audio data, comprising: receiving an encoded core audio signal, encoded radiation pattern metadata, and encoded audio object metadata; decoding the encoded core audio signal to determine a core audio signal; decoding the encoded radiation pattern metadata to determine a decoded radiation pattern; decoding the audio object metadata; and rendering the core audio signal based on the audio object metadata and the decoded radiation pattern. Method. 〔Aspect 9〕 The method according to aspect 8, wherein the audio object metadata includes at least one of time-varying 3 degrees of freedom (3DoF) or 6 degrees of freedom (6DoF) source orientation information. 〔Aspect 10〕 The method according to aspect 8 or 9, wherein the core audio signal includes a plurality of directional objects based on a cluster of objects, and the decoded radiation pattern represents a centroid reflecting an average value for each frequency band. 〔Aspect 11〕 The method according to any one of aspects 8 to 10, wherein the rendering is based at least in part on applying subband gains to the decoded core audio signal based on the decoded radiation data. 〔Aspect 12〕 The method according to any one of aspects 8 to 11, wherein the encoded radiation pattern metadata corresponds to a set that varies with time and frequency of spherical harmonic function coefficients. [Aspect 13] The method according to any one of aspects 8 to 12, wherein the encoded radiation pattern metadata includes audio object type metadata. [Aspect 14] The method according to aspect 13, wherein the audio object type metadata indicates parametric directivity pattern data, and the parametric directivity pattern data includes one or more functions selected from a list of functions consisting of a cosine function, a sine function, or a cardioid function. [Aspect 15] The method according to aspect 13, wherein the audio object type metadata indicates database directivity pattern data, and determining the decoded radiation pattern by decoding the encoded radiation pattern metadata includes querying a directivity data structure including the audio object type and corresponding directivity pattern data. [Aspect 16] The method according to aspect 13, wherein the audio object type metadata indicates dynamic directivity pattern data, and the dynamic directivity pattern data corresponds to a set that varies with time and frequency of spherical harmonic function coefficients. [Aspect 17] The method according to aspect 16, further comprising receiving the dynamic directivity pattern data before receiving the encoded core audio signal. [Aspect 18] An interface system; and An audio decoding apparatus having a control system, wherein the control system is: Receiving audio data corresponding to at least one audio object via the interface system, wherein the audio data includes a monophonic audio signal, audio object position metadata, audio object size metadata, and rendering parameters; Determining whether the rendering parameters indicate a position mode or a directivity mode; and when it is determined that the rendering parameters indicate a directivity mode, rendering the audio data for reproduction via at least one loudspeaker according to a directivity pattern indicated by at least one of the position metadata or the size metadata. Device. 〔Aspect 19〕 The device according to aspect 18, wherein rendering the audio data includes interpreting the audio object position metadata as audio object orientation metadata. 〔Aspect 20〕 The device according to aspect 19, wherein the audio object position metadata includes at least one of x, y, z coordinate data, spherical coordinate data, or cylindrical coordinate data, and the audio object orientation metadata includes yaw, pitch, and roll data. 〔Aspect 21〕 The device according to any one of aspects 18 to 20, wherein rendering the audio data includes interpreting the audio object size metadata as directivity metadata corresponding to the directivity pattern. 〔Aspect 22〕 The device according to any one of aspects 18 to 21, wherein rendering the audio data includes querying a data structure including a plurality of directivity patterns and mapping at least one of the position metadata or the size metadata to one or more of the directivity patterns. 〔Aspect 23〕 The apparatus according to aspect 22, wherein the control system is configured to receive the data structure via the interface system. 〔Aspect 24〕 The apparatus according to aspect 23, wherein the data structure is received prior to the audio data. 〔Aspect 25〕 The apparatus according to any one of aspects 18 to 24, wherein the audio data is received in a Dolby Atmos format. 〔Aspect 26〕 The apparatus according to any one of aspects 18 to 25, wherein the audio object position metadata corresponds to world coordinates or model coordinates.
Claims
Claim 1 A method for decoding audio data, comprising: receiving an encoded core audio signal, encoded radiation pattern metadata, and encoded audio object metadata, wherein the audio object metadata includes 6DoF source orientation information; decoding the encoded core audio signal to determine a core audio signal; decoding the encoded radiation pattern metadata to determine a decoded radiation pattern; rendering the core audio signal based on the audio object metadata and the decoded radiation pattern. The method. Claim 2 The method according to claim 1, wherein the core audio signal includes a plurality of directional objects based on a cluster of objects, and the decoded radiation pattern represents a centroid reflecting an average value for each frequency band. Claim 3 The method according to claim 1, wherein the encoded radiation pattern metadata corresponds to a set that varies over time and frequency of spherical harmonic function coefficients. Claim 4 The method according to claim 1, wherein the encoded radiation pattern metadata includes audio object type metadata. Claim 5 The audio object type metadata indicates parametric directivity pattern data, and the parametric directivity pattern data includes one or more functions selected from a list of functions consisting of a cosine function, a sine function, or a cardioid function. The method according to claim 4. Claim 6 The audio object type metadata indicates dynamic directivity pattern data, and the dynamic directivity pattern data corresponds to a set that varies over time and frequency of spherical harmonic function coefficients. The method according to claim 4. Claim 7 The method according to claim 6, further comprising receiving the dynamic directivity pattern data before receiving the encoded core audio signal. Claim 8 The method according to claim 1, wherein the rendering is based at least in part on applying a sub-band gain to the decoded core audio signal based on the decoded radiation pattern. Claim 9 The method according to claim 4, wherein the audio object type metadata indicates database-oriented pattern data, and determining the decoded radiation pattern by decoding the encoded radiation pattern metadata includes querying a directivity data structure including an audio object type and corresponding directivity pattern data.
10. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to claim 1.
11. An interface system; and An audio decoding apparatus having a control system, wherein the control system: Receiving audio data corresponding to at least one audio object via the interface system, the audio data including a monophonic audio signal, audio object position metadata, and rendering parameters, the audio object position metadata including 6DoF source orientation information; Determining whether the rendering parameter indicates a position mode or a directivity mode; and when it is determined that the rendering parameter indicates a directivity mode, rendering the audio data for reproduction via at least one loudspeaker according to a directivity pattern indicated by at least one of the audio object position metadata. Apparatus.
Citation Information
Patent Citations
Sound signal processing method and sound field reproduction system
JP2007043334A
Obtaining sparseness information for higher-order ambisonic audio renderers
JP2017520177A
Method, apparatus and system for encoding and decoding directional sound sources
JP2021518923A
Scalable downmix design with feedback for object-based surround codec
US20140023196A1
Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding
WO2019068638A1