encoding a scaled spatial component
By combining spatial audio coding and psychoacoustic audio coding technologies to identify the foreground and background components of stereo reverberation audio data, the problem of low coding efficiency in existing technologies is solved, and more efficient audio data compression and transmission are achieved.
Patent Information
- Application Number
- CN202080044605.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2020-06-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2040-06-23
AI Technical Summary
Existing psychoacoustic audio coding technologies have difficulty in effectively utilizing the masking effect of the human auditory system for efficient compression when processing ambisonic reverberation audio data, resulting in low coding efficiency.
A method combining spatial audio coding and psychoacoustic audio coding is adopted. The foreground and background components of the stereo reverberation coefficients are identified by a spatial audio coding device, and bit allocation and quantization are performed based on the psychoacoustic model to generate an efficient bit stream.
Improves the encoding efficiency of ambisonic audio data, reducing data transmission and storage requirements while maintaining audio quality.
Smart Images

Figure CN114008704B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 16 / 907,969, filed on June 22, 2020, entitled “CODING SCALED SPATIAL COMPONENTS,” which claims the benefit of U.S. Provisional Application No. 62 / 865,858, filed on June 24, 2019, entitled “CODINGSCALED SPATIAL COMPONENTS,” the entire contents of which are incorporated herein by reference in their entirety as if set forth in this disclosure. Technical Field
[0002] The present disclosure relates to audio data, and more particularly, to the coding of audio data. Background Art
[0003] Psychoacoustic audio coding refers to a process in which psychoacoustic models are used to compress audio data. Psychoacoustic audio codecs can exploit limitations in the human auditory system to compress audio data, taking into account limitations that occur due to spatial masking (e.g., two audio sources at the same location where one of the auditory sources masks the other in loudness), temporal masking (e.g., where one audio source masks the other in loudness), etc. Psychoacoustic models can attempt to model the human auditory system to identify masked or other portions of a sound field that are redundant, masked, or otherwise not perceivable by the human auditory system. Psychoacoustic audio codecs can also perform lossless compression by entropy encoding the audio data. Summary of the Invention
[0004] Generally, techniques for encoding and decoding scaled spatial components are described.
[0005] In one example, various aspects of the technology relate to a device configured to encode scene-based audio data, the device comprising: a memory configured to store the scene-based audio data; and one or more processors configured to: perform spatial audio encoding relative to the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; perform psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; determine a bit allocation relative to the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; scale the spatial components based on the bit allocation to the foreground audio signal to obtain scaled spatial components; quantize the scaled spatial components to obtain quantized spatial components; and specify the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0006] In another example, various aspects of the technology relate to a method for encoding scene-based audio data, the method comprising: performing spatial audio encoding relative to the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; performing psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; determining a bit allocation for the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; scaling the spatial components based on the bit allocation for the foreground audio signal to obtain scaled spatial components; quantizing the scaled spatial components to obtain quantized spatial components; and specifying the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0007] In another example, various aspects of the technology relate to a device configured to encode scene-based audio data, the device comprising: a component for performing spatial audio encoding relative to the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; a component for performing psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; a component for determining a bit allocation to the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; a component for scaling the spatial components based on the bit allocation to the foreground audio signal to obtain scaled spatial components; a component for quantizing the scaled spatial components to obtain quantized spatial components; and a component for specifying the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0008] In another example, various aspects of the technology relate to a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: perform spatial audio encoding relative to scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; perform psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; determine a bit allocation for the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; scale the spatial components based on the bit allocation for the foreground audio signal to obtain scaled spatial components; quantize the scaled spatial components to obtain quantized spatial components; and specify the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0009] In another example, various aspects of the technology relate to a device configured to decode a bitstream representing encoded scene-based audio data, the device comprising: a memory configured to store the bitstream, the bitstream comprising an encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of the encoded foreground audio signal; and one or more processors configured to: perform psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain a foreground audio signal; determine a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; dequantize the quantized spatial components to obtain scaled spatial components; descale the scaled spatial components based on the bit allocation for the encoded foreground audio signal to obtain spatial components; and reconstruct the scene-based audio data based on the foreground audio signal and the spatial components.
[0010] In another example, various aspects of the technology relate to a method for decoding a bitstream representing scene-based audio data, the method comprising: obtaining an encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of the encoded foreground audio signal from the bitstream; performing psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain the foreground audio signal; determining a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; dequantizing the quantized spatial components to obtain scaled spatial components; descaling the scaled spatial components based on the bit allocation to the encoded foreground audio signal to obtain spatial components; and reconstructing the scene-based audio data based on the foreground audio signal and the spatial components.
[0011] In another example, various aspects of the technology relate to a device configured to decode a bitstream representing encoded scene-based audio data, the device comprising: a component for obtaining an encoded foreground audio signal and corresponding quantized spatial components that define spatial characteristics of the encoded foreground audio signal from the bitstream; a component for performing psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain the foreground audio signal; a component for determining a bit allocation to the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; a component for dequantizing the quantized spatial components to obtain scaled spatial components; a component for descaling the scaled spatial components based on the bit allocation to the encoded foreground audio signal to obtain spatial components; and a component for reconstructing scene-based audio data based on the foreground audio signal and the spatial components.
[0012] In another example, various aspects of the technology relate to a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: obtain an encoded foreground audio signal and corresponding quantized spatial components that define spatial characteristics of the encoded foreground audio signal from a bitstream; perform psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain a foreground audio signal; determine a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; dequantize the quantized spatial components to obtain scaled spatial components; descale the scaled spatial components based on the bit allocation to the encoded foreground audio signal to obtain spatial components; and reconstruct scene-based audio data based on the foreground audio signal and the spatial components.
[0013] The details of one or more aspects of the technology are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the technology will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a diagram illustrating a system that can perform various aspects of the techniques described in this disclosure.
[0015] Figure 2 is a diagram illustrating another example of a system that can perform various aspects of the techniques described in this disclosure.
[0016] Figures 3A to 3C are illustrated in more detail in Figure 1 and Figure 2 A block diagram of an example of a psychoacoustic audio encoding device is shown in the example of FIG.
[0017] Figures 4A to 4C are illustrated in more detail in Figure 1 and Figure 2 A block diagram of an example of a psychoacoustic audio decoding device is shown in the example of FIG.
[0018] Figure 5 is shown in more detail in Figures 3A to 3C A block diagram of an example of an encoder is shown in the example of .
[0019] Figure 6 A more detailed diagram Figures 4A to 4C Block diagram of an example of a decoder.
[0020] Figure 7 is shown in more detail in Figures 3A to 3C A block diagram of an example of an encoder is shown in the example of .
[0021] Figure 8 is shown in more detail in Figures 4A to 4C00140 - 00141 - 00142 - 00143 ...
[0022] Figure 9A and Figure 9B is shown in more detail in Figures 3A to 3C Block diagram of another example of an encoder shown in the example of .
[0023] Figure 10A and Figure 10B is shown in more detail in Figures 4A to 4C Block diagram of another example of a decoder shown in the example of .
[0024] Figure 11 is a diagram illustrating an example of top-down quantization.
[0025] Figure 12 is a diagram illustrating an example of bottom-up quantization.
[0026] Figure 13 It is shown in the figure Figure 2 00105] A block diagram of example components of a source device is shown in the examples of FIG.
[0027] Figure 14 It is shown in the figure Figure 2 is a block diagram of exemplary components of a terminal device shown in the example of .
[0028] Figure 15 It is shown in the figure Figure 1 Flowchart of example operations of an audio encoder shown in an example of FIG. 1 in performing various aspects of the techniques described in this disclosure.
[0029] Figure 16 It is shown in the figure Figure 1 Flowchart of example operations of an audio decoder shown in the example of FIG. 1 in performing various aspects of the techniques described in this disclosure. DETAILED DESCRIPTION
[0030] There are different types of audio formats, including channel-based, object-based, and scene-based. Scene-based formats can use ambisonics technology. Ambisonics technology allows the sound field to be represented using a layered collection of elements that can be rendered to speaker feeds for most speaker configurations.
[0031] An example of a hierarchical collection of elements is the spherical harmonic coefficient (SHC) collection. The following expression demonstrates the use of SHC to describe or represent a sound field:
[0032]
[0033] This expression shows that at time t, at any point in the sound field The pressure at p i Can be by SHC Uniquely express. Here c is the speed of sound (~343 m / s), is the reference point (or observation point), j n (·) is a spherical Bessel function of order n, and are the spherical harmonic basis functions of order n and suborder m (which may also be referred to as spherical basis functions). It can be appreciated that the terms in the square brackets are signals (i.e. ), which can be approximated by various time-frequency transforms (such as discrete Fourier transform (DFT), discrete cosine transform (DCT), or wavelet transform). Other examples of hierarchical sets include sets of wavelet transform coefficients and other sets of coefficients of multiresolution basis functions.
[0034] SHC It can be physically acquired (e.g., recorded) by various microphone array configurations, or alternatively, it can be derived from a channel-based or object-based description of the sound field (e.g., a pulse code modulation (PCM) audio object containing audio objects and metadata defining the location of the audio objects within the sound field). The SHC (also known as the ambisonics coefficient) represents scene-based audio, where the SHC can be input to an audio encoder to obtain encoded SHC that can facilitate more efficient transmission or storage. For example, a method involving (1+4) 2 (25, hence fourth order) The fourth order representation of the coefficients.
[0035] As described above, SHC can be derived from microphone recordings using a microphone array. Poletti, M., "Three-Dimensional Surround Sound Systems Based on Spherical Harmonics," Audio Engineering Society Journal, Vol. 53, No. 11, November 2005, pp. 1004-1025, describes various examples of how SHC can be derived from a microphone array.
[0036] To illustrate how SHC can be derived from an object-based description, consider the following equation: The coefficients of the sound field corresponding to a single audio object It can be expressed as:
[0037]
[0038] where i is is a spherical Hankel function of order n (second kind), and is the position of the object. Knowing the object source energy g(ω) as a function of frequency (e.g., using time- frequency analysis techniques, such as performing a fast Fourier transform on a PCM stream), allows us to convert each PCM object and corresponding position into SHC Furthermore, (since the above is a linear and orthogonal decomposition) the coefficients for each object can be shown to be additive. In this way, multiple PCM objects (where a PCM object is one example of an audio object) can be represented by coefficients (e.g., as a sum of coefficient vectors for individual objects). In essence, the coefficients contain information about the soundfield (pressure as a function of 3D coordinates), and the above representation is a conversion from individual objects to an overall soundfield representation in the vicinity of the observation point The latter figures are described below in the context of SHC-based audio coding.
[0039] Figure 1 is a diagram showing a system 10 that can perform various aspects of the techniques described in this disclosure. As shown in the example of Figure 1 System 10 includes a content creator device 12 and a content consumer 14. While described in the context of a content creator device 12 and a content consumer 14, these techniques can be implemented in any context in which SHC (which can also be referred to as ambisonic coefficients) or any other hierarchical representation of a soundfield is encoded to form a bitstream representing audio data.
[0040] Furthermore, as some examples, content creator system 12 can represent a system including one or more computing devices of any form capable of implementing the techniques described in this disclosure, including a handset (or cellular telephone, including a so-called “smartphone,” or, in other words, a mobile telephone or handset), a tablet computer, a laptop computer, a desktop computer, an extended reality (XR) device (which can refer to any one or more of a virtual reality, VR, device, an augmented reality, AR, device, a mixed reality, MR, device, etc.), a game system, a disc player, a receiver (such as an audio / visual, A / V, receiver), or a specialized hardware.
[0041] Likewise, as some examples, content consumer 14 can represent a computing device of any form capable of implementing the techniques described in this disclosure, including a handset (or cellular telephone, including a so-called “smartphone,” or, in other words, a mobile handset or telephone), an XR device, a tablet computer, a television (including a so-called “smart television”), a set-top box, a laptop computer, a game system or console, a watch (including a so-called smart watch), wireless earphones (including so-called “smart earphones”), or a desktop computer.
[0042] Content creator system 12 may represent any entity that can generate audio content, and possibly video content, for consumption by content consumers, such as content consumer 14. Content creator system 12 may capture live audio data at an event, such as a sporting event, while also inserting various other types of additional audio data into the live audio content, such as commentary audio data, commercial audio data, intro or exit audio data, and the like.
[0043] Content consumer 14 represents an individual who owns or has access to an audio playback system 16, which may refer to any form of audio playback system capable of rendering high-order ambisonic audio data (which includes high-order audio coefficients, also known as spherical harmonic function coefficients) to speakers for playback as audio content. Figure 1 In the example of FIG, content consumer 14 includes audio playback system 16.
[0044] Ambisonic audio data may be defined in the spherical harmonics domain and rendered or otherwise transformed from the spherical harmonics function domain to the spatial domain, thereby producing audio content in the form of one or more speaker feeds. Ambisonic audio data may represent an example of "scene-based audio data" that describes an audio scene using ambisonic coefficients. Scene-based audio data differs from object-based audio data in that the entire scene is described (in the spherical harmonics domain), as opposed to discrete objects (in the spatial domain) as is common in object-based audio data. Scene-based audio data differs from channel-based audio data in that, as opposed to the spatial domain of channel-based audio data, scene-based audio data resides in the spherical harmonics domain.
[0045] In any case, the content creator system 12 includes microphones 18 that record or otherwise obtain live recordings in various formats, including directly as ambisonics coefficients and audio objects. When the microphone array 18 (which may also be referred to as "microphones 18") obtains live audio directly as ambisonics coefficients, the microphones 18 may include a transcoder, such as Figure 1 The ambisonic transcoder 20 is shown in the example of FIG.
[0046] In other words, although shown as separate from the microphones 5, a separate instance of the ambisonic transcoder 20 may be included within each microphone 5 to transcode the captured feed into the ambisonic coefficients 21. However, when not included within the microphone 18, the ambisonic transcoder 20 may transcode the live feed output from the microphone 18 into the ambisonic coefficients 21. In this regard, the ambisonic transcoder 20 may represent a unit configured to transcode microphone feeds and / or audio objects into the ambisonic coefficients 21. Thus, the content creator system 12 includes the ambisonic transcoder 20, such as integrated with the microphones 18, such as separate from the microphones 18, or some combination thereof.
[0047] The content creator system 12 may also include an audio encoder 22 configured to compress the ambisonic reverberation coefficients 21 to obtain a bitstream 31. The audio encoder 22 may include a spatial audio encoding device 24 and a psychoacoustic audio encoding device 26. The spatial audio encoding device 24 may represent a device capable of performing compression with respect to the ambisonic reverberation coefficients 21 to obtain intermediate formatted audio data 25 (which may also be referred to as "mezzanine formatted audio data 25" when the content creator system 12 represents a broadcast network as described in more detail below). The intermediate formatted audio data 25 may represent audio data that has been compressed using spatial audio compression but has not yet undergone psychoacoustic audio coding (e.g., AptX or Advanced Audio Codec - AAC, or other similar types of psychoacoustic audio coding, including various enhanced AAC (EAAC) variants such as High Efficiency AAC (HE-AAC, HE-AAC v2 (also known as eAAC+), etc.).
[0048] The spatial audio encoding device 24 may be configured to compress the ambisonic reverberation coefficients 21. That is, the spatial audio encoding device 24 may compress the ambisonic reverberation coefficients 21 using a decomposition involving application of a linear reversible transform (LIT). One example of a linear reversible transform is called "singular value decomposition" ("SVD"), principal component analysis ("PCA"), or eigenvalue decomposition, which may represent different examples of linear reversible decomposition.
[0049] In this example, the spatial audio encoding device 24 may apply SVD to the ambisonic reverberation coefficients 21 to determine a decomposed version of the ambisonic reverberation coefficients 21. The decomposed version of the ambisonic reverberation coefficients 21 may include one or more of the dominant audio signals and one or more corresponding spatial components that describe the spatial characteristics (e.g., direction, shape, and width) of the associated dominant audio signals. Thus, the spatial audio encoding device 24 may apply the decomposition to the ambisonic reverberation coefficients 21 to decouple the energy (as represented by the dominant audio signals) from the spatial characteristics (as represented by the spatial components).
[0050] The spatial audio encoding device 24 may analyze the decomposed version of the ambisonic reverberation coefficients 21 to identify various parameters that may facilitate reordering the decomposed version of the ambisonic reverberation coefficients 21. The spatial audio encoding device 24 may reorder the decomposed version of the ambisonic reverberation coefficients 21 based on the identified parameters, wherein the reordering may improve encoding and decoding efficiency, assuming that the transform may reorder the ambisonic reverberation coefficients across a frame of the ambisonic reverberation coefficients (where a frame typically contains M samples of the decomposed version of the ambisonic reverberation coefficients 21, and in some examples, M is set to 1024).
[0051] After reordering the decomposed versions of the ambisonic reverberation coefficients 21, the spatial audio encoding device 24 may select one or more of the decomposed versions of the ambisonic reverberation coefficients 21 as a representation of a foreground (or, in other words, distinct, dominant, or significant) component of the sound field. The spatial audio encoding device 24 may specify a decomposed version of the ambisonic reverberation coefficients 21 that represents the foreground component (which may also be referred to as a "dominant sound signal," "dominant audio signal," or "dominant sound component") and associated directional information (which may also be referred to as a "spatial component," or in some cases, a so-called "V-vector" identifying the spatial characteristics of the corresponding audio object). The spatial component may represent a vector having multiple different elements (which may be referred to as "coefficients" in vector terms) and may thus be referred to as a "multidimensional vector."
[0052] The spatial audio encoding device 24 may then perform a sound field analysis on the ambisonic reverberation coefficients 21 to at least partially identify the ambisonic reverberation coefficients 21 that represent one or more background (or, in other words, ambient) components of the sound field. Background components may also be referred to as "background audio signals" or "ambient audio signals." In some examples, assuming that the background audio signal may only include a subset of any given sample of the ambisonic reverberation coefficients 21 (e.g., samples corresponding to zero- and first-order spherical basis functions, but not samples corresponding to second- or higher-order spherical basis functions), the spatial audio encoding device 24 may perform energy compensation on the background audio signal. In other words, when performing order reduction, the spatial audio encoding device 24 may amplify (e.g., add energy to / subtract energy from) the remaining background ambisonic reverberation coefficients of the ambisonic reverberation coefficients 21 to compensate for the change in total energy caused by performing the order reduction.
[0053] The spatial audio encoding device 24 may then perform a form of interpolation (which is another way of referencing spatial components) with respect to the foreground directional information, and then perform order reduction with respect to the interpolated foreground directional information to generate reduced-order foreground directional information. In some examples, the spatial audio encoding device 24 may further perform quantization with respect to the reduced-order foreground directional information, thereby outputting the encoded foreground directional information. In some cases, this quantization may include scalar / entropy quantization, which may be in the form of vector quantization. The spatial audio encoding device 24 may then output intermediate formatted audio data 25 as a background audio signal, a foreground audio signal, and the quantized foreground directional information.
[0054] In any case, in some examples, the background audio signal and the foreground audio signal may include a transmission channel. That is, the spatial audio encoding device 24 may output a transmission channel for each frame including the ambisonic reverberation coefficient 21 of a corresponding one of the background audio signals (e.g., M samples of one of the ambisonic reverberation coefficients 21 corresponding to a zero-order or first-order spherical basis function) and for each frame of the foreground audio signal (e.g., M samples of an audio object decomposed from the ambisonic reverberation coefficient 21). The spatial audio encoding device 24 may further output side information (which may also be referred to as "sideband information") including quantized spatial components corresponding to each of the foreground audio signals.
[0055] Together, the transmission channel and side information can be Figure 1 In the example of FIG, the audio data 25 is represented as Ambisonics Transport Format (ATF) audio data 25 (which is another way of referring to intermediate formatted audio data). In other words, the AFT audio data 25 may include transmission channels and side information (which may also be referred to as "metadata"). As an example, the ATF audio data 25 may conform to the HOA (Higher Order Ambisonics) Transport Format (HTF). More information about HTF can be found in the European Telecommunications Standards Institute (ETSI) Technical Specification (TS) ETSI TS 103 589 V1.1.1, dated June 2018 (2018-06), entitled "Higher Order Ambisonics (HOA) Transport Format." Therefore, the ATF audio data 25 may be referred to as HTF audio data 25.
[0056] The spatial audio encoding device 24 may then send or otherwise output the ATF audio data 25 to a psychoacoustic audio encoding device 26. The psychoacoustic audio encoding device 26 may perform psychoacoustic audio encoding with respect to the ATF audio data 25 to generate a bitstream 31. The psychoacoustic audio encoding device 26 may operate according to a standardized, open source, or proprietary audio codec process. For example, the psychoacoustic audio encoding device 26 may operate according to, for example, the Unified Speech and Audio Codec denoted as "USAC" as formulated by the Moving Picture Experts Group (MPEG), the MPEG-H 3D Audio Codec standard, the MPEG-I Immersive Audio standard, or a proprietary standard such as AptX. TM The content creator system 12 may then transmit the bitstream 31 to the content consumer 14 via a transmission channel.
[0057] In some examples, the psychoacoustic audio encoding device 26 may represent one or more instances of a psychoacoustic audio codec, each of which is used to encode a transmission channel of the ATF audio data 25. In some cases, the psychoacoustic audio encoding device 26 may represent one or more instances of an AptX encoding unit (as described above). In some cases, the psychoacoustic audio codec unit 26 may invoke an instance of a stereo encoding unit for each transmission channel of the ATF audio data 25.
[0058] In some examples, to generate different representations of the sound field using the ambisonics coefficients (which is again an example of audio data 21), the audio encoder 22 may use a codec scheme for a ambisonics representation of the sound field, which is referred to as mixed-order ambisonics (MOA), as discussed in more detail in U.S. application serial number 15 / 672,058, filed on August 8, 2017, and U.S. patent publication number 2019 / 0007781, published on January 3, 2019, entitled “MIXED-ORDERAMBISONICS (MOA) AUDIO DATA FOR COMPUTER-MEDIATED REALITY SYSTEMS.”
[0059] To generate a particular MOA representation of a sound field, the audio encoder 22 may generate a partial subset of the complete set of ambisonic coefficients. For example, each MOA representation generated by the audio encoder 22 may provide accuracy with respect to some areas of the sound field, but less accuracy in other areas. In one example, the MOA representation of the sound field may include eight (8) of the ambisonic coefficients as uncompressed ambisonic coefficients, while a third-order ambisonic representation of the same sound field may include sixteen (16) of the ambisonic coefficients as uncompressed ambisonic coefficients. Thus, each MOA representation of the sound field generated as a partial subset of the ambisonic coefficients may be less storage intensive and less bandwidth intensive (if and when transmitted as part of the bitstream 31 over the illustrated transmission channel) than a corresponding third-order ambisonic representation of the same sound field generated from the ambisonic coefficients.
[0060] Although described with respect to an MOA representation, the techniques of this disclosure may also be performed with respect to a full-order ambisonic (FOA) representation, where all ambisonic coefficients for a given order N are used to represent the sound field. In other words, rather than using a partially non-zero subset of the ambisonic coefficients to represent the sound field, the sound field representation generator 302 may use all ambisonic coefficients for a given order N to represent the sound field, such that the total number of ambisonic coefficients is equal to (N+1). 2 .
[0061] In this regard, higher-order ambisonics audio data (which is another way of referring to ambisonics coefficients in MOA representation or FOA representation) may include higher-order ambisonics coefficients associated with spherical basis functions having an order of one or less (which may be referred to as “first-order ambisonics audio data”), higher-order ambisonics coefficients associated with spherical basis functions having mixed orders and sub-orders (which may be referred to as the “MOA representation” discussed above), or higher-order ambisonics coefficients associated with spherical basis functions having an order greater than one (which is referred to as the “FOA representation” above).
[0062] In addition, although Figure 1 31 is shown as being sent directly to content consumer 14, but content creator system 12 may output bitstream 31 to an intermediary device located between content creator system 12 and content consumer 14. The intermediary device may store bitstream 31 for later delivery to content consumer 14, which may request the bitstream. The intermediary device may include a file server, web server, desktop computer, laptop computer, tablet computer, mobile phone, smartphone, or any other device capable of storing bitstream 31 for later retrieval by an audio decoder. The intermediary device may reside in a content delivery network capable of streaming bitstream 31 (and possibly in conjunction with sending a corresponding video data bitstream) to subscribers (e.g., content consumers 14) requesting the bitstream 31.
[0063] Alternatively, the content creator system 12 may store the bitstream 31 to a storage medium such as a compact disc, a digital video disc, a high-definition video disc, or other storage medium, most of which are capable of being read by a computer and thus may be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this context, the distribution channels may refer to those channels through which the content stored to these media is distributed (and may include retail stores and other store-based delivery mechanisms). In any case, the technology of this disclosure should therefore not be limited in this regard. Figure 1 .
[0064] like Figure 1 As further shown in the example of , content consumer 14 includes an audio playback system 16. Audio playback system 16 may represent any audio playback system capable of playing back multi-channel audio data. Audio playback system 16 may further include an audio decoding device 32. Audio decoding device 32 may represent a device configured to decode ambisonic reverberation coefficients 11′ from a bitstream 31, where ambisonic reverberation coefficients 11′ may be similar to ambisonic reverberation coefficients 11, but differ due to lossy operations (e.g., quantization) and / or transmission via a transmission channel.
[0065] The audio decoding device 32 may include a psychoacoustic audio decoding device 34 and a spatial audio decoding device 36. The psychoacoustic audio decoding device 34 may represent a unit configured to operate reciprocally with the psychoacoustic audio encoding device 26 to reconstruct the ATF audio data 25′ from the bitstream 31. Likewise, the prime relative to the ATF audio data 25 output from the psychoacoustic audio decoding device 34 indicates that the ATF audio data 25′ may differ slightly from the ATF audio data 25 due to lossy or other operations performed during compression of the ATF audio data 25. The psychoacoustic audio decoding device 34 may be configured to perform decompression according to a standardized, open source, or proprietary audio codec process (e.g., the aforementioned AptX, a variant of AptX, AAC, a variant of AAC, etc.).
[0066] Although the following description is primarily about AptX, the techniques can be applied with respect to other psychoacoustic audio codecs. Examples of other psychoacoustic audio codecs include Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus, and Windows Media Audio (WMA).
[0067] In any case, psychoacoustic audio decoding device 34 can perform psychoacoustic decoding with respect to foreground audio objects specified in bitstream 31 and encoded ambisonic coefficients representing background audio signals specified in bitstream 31. In this way, psychoacoustic audio decoding device 34 can obtain ATF audio data 25' and output ATF audio data 25' to spatial audio decoding device 36.
[0068] Spatial audio decoding device 36 can represent a unit configured to operate inversely to spatial audio encoding device 24. That is, spatial audio decoding device 36 can dequantize foreground direction information specified in bitstream 31. Spatial audio decoding device 36 can further perform dequantization with respect to the dequantized foreground direction information to obtain decoded foreground direction information. Spatial audio decoding device 36 can next perform interpolation with respect to the decoded foreground direction information, and then determine ambisonic coefficients representing foreground components based on the decoded foreground audio signals and the interpolated foreground direction information. Spatial audio decoding device 36 can then determine ambisonic coefficients 11' based on the determined ambisonic coefficients representing foreground audio signals and the decoded ambisonic coefficients representing background audio signals.
[0069] Audio playback system 16 can render ambisonic coefficients 11' to output speaker feeds 39 after decoding bitstream 31 to obtain ambisonic coefficients 11'. Audio playback system 16 can include a number of different audio renderers 38. Audio renderers 38 can each provide a different form of rendering, which can include one or more of various ways of performing vector-based amplitude panning (VBAP), one or more of various ways of performing binaural rendering (e.g., head-related transfer function - HRTF, binaural room impulse response - BRIR, etc.), and / or one or more of various ways of performing soundfield synthesis.
[0070] Audio playback system 16 can output speaker feeds 39 to one or more of speakers 40. Speaker feeds 39 can drive speakers 40. Speakers 40 can represent loudspeakers (e.g., transducers placed in a cabinet or other enclosure), earphone speakers, or any other type of transducer capable of emitting sound based on an electrical signal.
[0071] To select an appropriate renderer, or in some cases, generate an appropriate renderer, the audio playback system 16 may obtain loudspeaker information 41 indicating the number of speakers 40 and / or the spatial geometry of the speakers 40. In some cases, the audio playback system 16 may obtain the loudspeaker information 41 using a reference microphone and drive the speakers 40 in a manner that dynamically determines the loudspeaker information 41. In other examples, or in conjunction with the dynamic determination of the loudspeaker information 41, the audio playback system 16 may prompt a user to interface with the audio playback system 16 and enter the loudspeaker information 41.
[0072] The audio playback system 16 may select one of the audio renderers 38 based on the speaker information 41. In some cases, the audio playback system 16 may generate one of the audio renderers 38 based on the speaker information 41 when no audio renderer 38 is within a certain threshold similarity metric (in terms of loudspeaker geometry) of the similarity metric specified in the speaker information 41. In some cases, the audio playback system 16 may generate one of the audio renderers 38 based on the speaker information 41 without first attempting to select an existing one of the audio renderers 38.
[0073] Although described with respect to speaker feeds 39, the audio playback system 16 may render a headphone feed from the speaker feeds 39 or directly from the ambisonics coefficients 11', thereby outputting the headphone feed to the headphone speakers. The headphone feed may represent a binaural audio speaker feed that the audio playback system 16 renders using a binaural audio renderer. As described above, the audio encoder 22 may invoke the spatial audio encoding device 24 to perform spatial audio encoding (or otherwise compress) the ambisonics audio data 21 and thereby obtain ATF audio data 25. During the application of spatial audio encoding to the ambisonics audio data 21, the spatial audio encoding device 24 may obtain a foreground audio signal and a corresponding spatial component, which are designated as a transmission channel and accompanying metadata (or side information), respectively, in encoded form.
[0074] As described above, the spatial audio encoding device 24 may apply vector quantization with respect to the spatial components and prior to specifying the spatial components as metadata in the AFT audio data 25. The psychoacoustic audio encoding device 26 may quantize each of the transmission channels of the ATF audio data 25 independently of the quantization of the spatial components performed by the spatial audio encoding device 24. Because the spatial components provide spatial characteristics for the corresponding foreground audio signals, the independent quantization may result in different errors between the spatial components and the foreground audio signals, which may result in audio artifacts upon playback, such as incorrect positioning of the foreground audio signals within the reconstructed sound field, poor spatial resolution for higher-quality foreground audio signals, and other anomalies that may cause distracting or apparent inaccuracies during the reproduction of the sound field.
[0075] According to various aspects of the techniques described in this disclosure, the spatial audio encoding device 24 and the psychoacoustic audio encoding device 26 are integrated, and the psychoacoustic audio encoding device 26 can be combined with a spatial component quantizer (SCQ) 46 to offload quantization from the spatial audio encoding device 24. The SCQ 46 can scale the spatial components based on the bit allocation specified for the transmission channel, thereby reducing the dynamic range of the spatial components and thereby potentially reducing the degree of quantization applied to the spatial components. Reducing the degree of quantization can improve the spatial accuracy of the reconstructed HTF audio data 25' and thereby potentially reduce the injection of the audio artifacts described above, which can improve the operation of the system 10 itself.
[0076] In operation, the spatial audio encoding device 24 may perform spatial audio encoding on the scene-based audio data 21 to obtain a foreground audio signal and corresponding spatial components. However, the spatial audio encoding performed by the spatial audio encoding device 24 omits the aforementioned spatial component quantization, as the quantization is again offloaded to the psychoacoustic audio encoding device 26. The spatial audio encoding device 24 may output the ATF audio data 25 to the psychoacoustic audio encoding device 26.
[0077] The audio encoder 22 calls the psychoacoustic audio encoding device 26 to perform psychoacoustic audio encoding with respect to the foreground audio signal to obtain an encoded foreground audio signal. In some examples, the psychoacoustic audio encoding device 26 may perform psychoacoustic audio encoding according to a compression algorithm (including any of the various versions of AptX listed above). Figure 5 to Figure 1 The example of 0 generally describes the AptX compression algorithm.
[0078] The psychoacoustic audio encoding device 26 may determine a bit allocation for the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal. The psychoacoustic audio encoding device 26 may call the SCQ 46, thereby passing the bit allocation to the SCQ 46. The SCQ 46 may scale the spatial components based on the bit allocation of the foreground audio signal to obtain scaled spatial components. The SCQ 46 may then quantize (e.g., vector quantize) the scaled spatial components to obtain quantized spatial components. The psychoacoustic audio encoding device 26 may then specify the encoded foreground audio signal and the quantized spatial components in the bitstream 31.
[0079] As described above, the audio decoder 32 can operate inversely with the audio encoder 22. Thus, the audio decoder 32 can obtain the bitstream 31 and call the psychoacoustic audio decoding device 34 to perform psychoacoustic audio decoding on the encoded foreground audio signal to obtain the foreground audio signal. As described above, the psychoacoustic audio decoding device 34 can perform psychoacoustic audio decoding according to the AptX decompression algorithm. Again, below with respect to Figure 5 to Figure 10's example describes more information about the AptX decompression algorithm.
[0080] In any case, when psychoacoustic audio encoding is performed with respect to the foreground audio signal, the psychoacoustic audio decoding device 34 may determine a bit allocation for the encoded foreground audio signal. The psychoacoustic audio decoding device 34 may call the SCD 54, thereby communicating the bit allocation to the SCD 54. The SCD 54 may descale the scaled spatial components based on the bit allocation of the foreground audio signal to obtain quantized spatial components. The SCD 54 may then dequantize (e.g., vector dequantize) the scaled spatial components to obtain spatial components. The psychoacoustic audio decoding device 34 may reconstruct the ATF audio data 25' based on the foreground audio signal and the spatial components. The spatial audio decoding device 36 may then reconstruct the scene-based audio data 21' based on the foreground audio signal and the spatial components of the ATF audio data 25'.
[0081] Figure 2 is a diagram illustrating another example of a system that can perform various aspects of the techniques described in this disclosure. Figure 2 The system 110 may be represented in Figure 1 An example of system 10 is shown in the example of FIG. Figure 2 As shown in the example of , system 110 includes a source device 112 and a sink device 114, where source device 112 may represent an example of content creator system 12 and sink device 114 may represent an example of content consumer 14 and / or audio playback system 16.
[0082] Although described with respect to source device 112 and sink device 114, in some cases source device 112 may operate as a sink device, and in these and other cases sink device 114 may operate as a source device. Figure 2 The example of system 110 shown in FIG. 1 is merely one example illustrating various aspects of the techniques described in this disclosure.
[0083] In any case, as described above, source device 112 may represent any form of computing device capable of implementing the techniques described in this disclosure, including handheld computers (or cellular phones, including so-called "smart phones"), tablet computers, so-called smart phones, remotely piloted aircraft (such as so-called "drones"), robots, desktop computers, receivers (such as audio / visual - AV- receivers), set-top boxes, televisions (including so-called "smart TVs"), media players (such as digital video disc players, streaming media players, Blu-ray disc players), and other similar devices. TM Player, etc.), or any other device capable of transmitting audio data wirelessly to a terminal device via a personal area network (PAN). For illustration, it is assumed that the source device 12 represents a smart phone.
[0084] The terminal device 114 may represent any form of computing device capable of implementing the techniques described in this disclosure, including a handset (or in other words, a cellular phone, a mobile phone, a mobile handset, etc.), a tablet computer, a smart phone, a desktop computer, a wireless headset (which may include wireless headsets with or without a microphone, as well as so-called smart wireless headsets that include additional functionality such as fitness monitoring, onboard music storage and / or playback, dedicated cellular capabilities, etc.), a wireless speaker (including so-called "smart speakers"), a watch (including so-called "smart watches"), or any other device capable of reproducing a sound field based on audio data wirelessly transmitted via a PAN. And, for purposes of illustration, it is assumed that the terminal device 114 represents a wireless headset.
[0085] like Figure 2 As shown in the example of , source device 112 includes one or more application programs ("applications") 118A-118N ("applications 118"), a mixing unit 120, an audio encoder 122 (which includes a spatial audio encoding device (SAED) 124 and a psychoacoustic audio encoding device (PAED) 126), and a wireless connection manager 128. Figure 2 Not shown in the example of , source device 112 may include a number of other elements that support the operation of application 118, including an operating system, various hardware and / or software interfaces (such as user interfaces, including graphical user interfaces), one or more processors, memory, storage devices, etc.
[0086] Each application 118 represents software (such as a set of instructions stored to a non-transitory computer-readable medium) that, when executed by one or more processors of source device 112, configures system 110 to provide some functionality. To name a few examples, applications 118 may provide messaging functionality (such as access to email, text messaging, and / or video messaging), voice calling functionality, video conferencing functionality, calendaring functionality, audio streaming functionality, directions functionality, mapping functionality, and gaming functionality. Applications 118 may be first-party applications designed and developed by the same company that designs and sells the operating system executed by source device 112 (and are typically pre-installed on source device 112), or third-party applications accessible via a so-called "app store" or potentially pre-installed on source device 112. Each application 118, when executed, may output audio data 119A-119N ("audio data 119"), respectively.
[0087] In some examples, audio data 119 may be received from a microphone (not shown, but similar to a microphone) connected to source device 112. Figure 1 The audio data 119 may include similar Figure 1The ambisonic reverberation coefficients of the ambisonic reverberation audio data 21 discussed in the example of , wherein the ambisonic reverberation audio data may be referred to as “scene-based audio data.” Thus, the audio data 119 may also be referred to as “scene-based audio data 119” or “ambisonic reverberation audio data 119.”
[0088] Although described with respect to ambisonic audio data, the techniques may be performed with respect to ambisonic audio data that does not necessarily include coefficients corresponding to so-called "higher-order" spherical basis functions (e.g., spherical basis functions having an order greater than one). Thus, the techniques may be performed with respect to ambisonic audio data that includes coefficients corresponding to only the zeroth-order spherical basis functions, or only the zeroth-order and first-order spherical basis functions.
[0089] The mixing unit 120 represents a unit configured to mix one or more of the audio data 119 output by the application 118 (as well as other audio data output by the operating system, such as alarms or other tones, including keyboard press tones, ringtones, etc.) to generate mixed audio data 121. Audio mixing can refer to a process in which multiple sounds (as described in the audio data 119) are combined into one or more channels. During mixing, the mixing unit 120 can also manipulate and / or enhance the volume level (which can also be referred to as "gain level"), frequency content, and / or panoramic position of the ambisonic audio data 119. In the context of streaming the ambisonic audio data 119 over a wireless PAN session, the mixing unit 120 can output the mixed audio data 121 to an audio encoder 122.
[0090] The audio encoder 122 may be similar to (or otherwise substantially similar to) the audio encoder described above in Figure 1 That is, the audio encoder 122 may represent a unit configured to encode the mixed audio data 121 and thereby obtain encoded audio data in the form of a bitstream 131. In some examples, the audio encoder 122 may encode a single one of the audio data 119.
[0091] For illustration, referring to an example of PAN protocol, Many different types of audio codecs (a word created by combining the words "encode" and "decode") are provided, and can be expanded to include vendor-specific audio codecs. The Advanced Audio Distribution Profile (A2DP) of the [A2DP specification] indicates that support for A2DP requires support for the sub-band codecs specified in A2DP. A2DP also supports codecs described in MPEG-1 Part 3 (MP2), MPEG-2 Part 3 (MP3), MPEG-2 Part 7 (Advanced Audio Coding - AAC), MPEG-4 Part 3 (High Efficiency - AAC - HE-AAC), and Adaptive Transform Acoustic Coding (ATRAC). In addition, as mentioned above, A2DP supports vendor-specific codecs such as aptX TM and various other versions of aptX (e.g., Enhanced aptX - E-aptX, aptX live, and aptX High Definition - aptX-HD).
[0092] The audio encoder 122 may operate in accordance with any of the audio codecs listed above, as well as one or more of the audio codecs not listed above that operate to encode the mixed audio data 121 to obtain the encoded audio data 131 (which is another way of referring to the bitstream 131). The audio encoder 122 may first call the SAED 124, which may be similar to (or otherwise substantially similar to) Figure 1 The SAED 124 may perform the above-described spatial audio compression on the mixed audio data to obtain ATF audio data 125 (which may be similar to (or otherwise substantially similar to) Figure 1 1). The SAED 124 may output the ATF audio data 25 to the PAED 126.
[0093] PAED 126 may be similar to (or otherwise substantially similar to) Figure 1 2. The PAED 126 may perform psychoacoustic audio encoding according to any of the aforementioned codecs (including AptX and its variants) to obtain a bitstream 131. The audio encoder 122 may output the encoded audio data 131 to one of the wireless communication units 130 managed by the wireless connection manager 128 (e.g., wireless communication unit 130A).
[0094] The wireless connection manager 128 may represent a unit configured to allocate bandwidth within certain frequencies of the available spectrum to different wireless communication units in the wireless communication units 130. For example, The communication protocol operates in the 2.5 GHz spectrum range, which overlaps with the spectrum range used by various WLAN communication protocols. The wireless connection manager 128 can allocate a portion of the bandwidth to the wireless device during a given time. The wireless connection manager 128 may include multiple protocols and allocate different portions of bandwidth to overlapping WLAN protocols at different times. The allocation of bandwidth, etc., is defined by a scheme 129. The wireless connection manager 128 may expose various application programmer interfaces through which the allocation of bandwidth and other aspects of the communication protocols may be adjusted to achieve a specified quality of service (QoS). In other words, the wireless connection manager 128 may provide an API to adjust the scheme 129, which controls the operation of the wireless communication unit 130 to achieve a specified QoS.
[0095] In other words, the wireless connection manager 128 can manage the coexistence of multiple wireless communication units 130 operating within the same spectrum, such as certain WLAN communication protocols and some PAN protocols as described above. The wireless connection manager 128 can include a coexistence scheme 129 (in Figure 2 129 ), which indicates when (eg, intervals) each of the wireless communication units 130 may send packets, how many packets to send, the size of the packets sent, and the like.
[0096] The wireless communication units 130 may each represent a wireless communication unit 130 that operates according to one or more communication protocols to transmit a bitstream 131 to a terminal device 114 via a transmission channel. Figure 2 In the example of FIG, for illustration, it is assumed that the wireless communication unit 130A is based on It is further assumed that wireless communication unit 130A operates in accordance with A2DP to establish a PAN link (via a transmit channel) to allow the bitstream 131 to be delivered from source device 112 to sink device 114. Although described with respect to a PAN link, the present invention may be described with respect to a cellular connection (such as so-called 3G, 4G and / or 5G cellular data services), WiFi, or any other similar communication protocol suite. TM Various aspects of the technology can be implemented using any type of wired or wireless connection.
[0097] about More information about the communication protocol suite can be found in the document entitled "Bluetooth Core Specification v 5.0", published on December 6, 2016, and available at www.bluetooth.org / en-us / specification / adapted-specifications. More information about A2DP can be found in the document entitled "Advanced Audio Distribution Profile Specification", version 1.3.1, published on July 14, 2015.
[0098] The wireless communication unit 130A can output the bit stream 131 to the terminal device 114 via a transmission channel, which is assumed to be a wireless channel in the example of Bluetooth. Figure 2 14. Although shown in FIG. 14 as being sent directly to the end device 114, the source device 112 may output the bitstream 131 to an intermediate device positioned between the source device 112 and the end device 114. The intermediate device may store the bitstream 131 for later delivery to the end device 14 that may request the bitstream 131. The intermediate device may include a file server, a web server, a desktop computer, a laptop computer, a tablet computer, a mobile phone, a smartphone, or any other device capable of storing the bitstream 31 for later retrieval by an audio decoder. The intermediate device may reside in a content delivery network that is capable of streaming the bitstream 31 (and possibly in conjunction with sending a corresponding video data bitstream) to subscribers (e.g., content consumers 14) requesting the bitstream 31.
[0099] Alternatively, source device 112 may store bitstream 31 to a storage medium, such as a compact disc, digital video disc, high-definition video disc, or other storage medium, most of which are capable of being read by a computer and thus may be referred to as computer-readable storage medium or non-transitory computer-readable storage medium. In this context, the distribution channels may refer to those channels through which the content stored to these media is distributed (and may include retail stores and other store-based delivery mechanisms). In any case, the technology of this disclosure should therefore not be limited in this regard. Figure 2 .
[0100] like Figure 2 As further shown in the example of FIG, the terminal device 114 includes a wireless connection manager 150 that manages one or more of wireless communication units 152A-152N ("wireless communication units 152") according to a scheme 151, an audio decoder 132 (including a psychoacoustic audio decoding device - PADD-134 and a spatial audio decoding device - SADD-136), and one or more speakers 140A-140N ("speakers 140", which may be similar to Figure 1 40 shown in the example of FIG. 1 ). Wireless connection manager 150 may operate in a manner similar to that described above with respect to wireless connection manager 128, exposing an API to coordinate a scheme 151 by which operation of wireless communication unit 152 achieves a specified QoS.
[0101] The wireless communication unit 152 may be similar in operation to the wireless communication unit 130, except that the wireless communication unit 152 operates reciprocally with the wireless communication unit 130 to receive the bit stream 131 via the transmission channel. Assume that one of the wireless communication units 152 (e.g., the wireless communication unit 152A) receives the bit stream 131 according to The wireless communication unit 152A may output the bit stream 131 to the audio decoder 132.
[0102] The audio decoder 132 may operate in a reciprocal manner to the audio encoder 122. The audio decoder 132 may operate in concert with one or more of the audio codecs listed above, as well as any of the audio codecs not listed above that operate to decode the encoded audio data 131 to obtain the mixed audio data 121′. Again, the prime notation with respect to the “mixed audio data 121” indicates that there may be some loss due to quantization or other lossy operations that occurred during encoding by the audio encoder 122.
[0103] The audio decoder 132 may call the PADD 134 to perform psychoacoustic audio decoding on the bitstream 131 to obtain the ATF audio data 125', and the PADD 134 may output the ATF audio data 125' to the SADD 136. The SADD 136 may perform spatial audio decoding to obtain the mixed audio data 121'. Although for ease of illustration, Figure 2 The renderer is not shown in the example (similar to Figure 1 38), but the audio decoder 132 may render the mixed audio data 121 ' to the speaker feed (using any one of the renderers, such as described above with respect to Figure 1 The renderer 38 discussed in the example of ) and the speaker feed is output to one or more of the speakers 140.
[0104] Each of the speakers 140 represents a transducer configured to be fed by the speaker to reproduce the sound field. The transducer may be integrated into Figure 2 14. The example of the ...
[0105] As described above, PAED 126 may perform various aspects of the quantization techniques described above with respect to PAED 26 to quantize spatial components based on their foreground audio signal-dependent bit allocations. PADD 134 may also perform various aspects of the quantization techniques described above with respect to PADD 34 to dequantize quantized spatial components based on their foreground audio signal-dependent bit allocations. Figure 3A and Figure 3BThe example provides more information about PAED 126, while the Figure 4A and Figure 4B The examples in
[00145] provide more information about PADD 134.
[0106] Figures 3A to 3C are illustrated in more detail in Figure 1 and Figure 2 A block diagram of an example of a psychoacoustic audio coding device is shown in the example of FIG. Figure 3A , the psychoacoustic audio encoder 226A may represent an example of the PADD 26 and / or the PADD 126. The PADD 226A may receive the transport channels 225A-225N from the AFT encoder 224 (where ATF encoder may represent another way to refer to the spatial audio encoding device 24). As described above with respect to the spatial audio encoding device 24, the ATF encoder 224 may perform spatial audio encoding with respect to the ambisonic reverberation coefficients 221 (which may represent an example of the ambisonic reverberation coefficients 21).
[0107] PADD 226A may invoke instances of stereo encoders 250A-250N ("stereo encoders 250"), which, as discussed in greater detail below, may perform psychoacoustic audio coding according to a stereo compression algorithm. Stereo encoders 250 may each process two transport channels to generate sub-bitstreams 233A-233N ("sub-bitstreams 233").
[0108] To compress the transmission channels, the stereo encoder 250 may perform a shape and gain analysis with respect to each of the transmission channels 225 to obtain a shape and a gain representing the transmission channel 225. The stereo encoder 250 may also predict a first transmission channel of a pair of transmission channels 225 from a second transmission channel of the pair of transmission channels 225, thereby predicting a gain and a shape representing the first transmission channel from the gain and the shape representing the second transmission channel to obtain a residual.
[0109] Before performing separate prediction of the gain, the stereo encoder 250 may first perform quantization relative to the gain of the second transmission channel to obtain a coarse quantized gain and one or more fine quantized residuals. Additionally, the stereo encoder 250 may perform quantization (e.g., vector quantization) relative to the shape of the second transmission channel to obtain a quantized shape before performing separate prediction of the shape. The stereo encoder 250 may then use the quantized coarse and fine energies and the quantized shape from the second transmission channel to predict the first transmission channel from the second transmission channel to predict the quantized coarse and fine energies and the quantized shape from the first transmission channel.
[0110] When quantizing the transmission channel, stereo encoder 250 may determine bit allocations 251A-251N for energy and shape ("bit allocations 251"), which indicate the number of bits used to represent each of the quantized coarse and fine energies and each of the quantized shapes. Stereo encoder 250 may output bit allocations 251 to SCQ 46.
[0111] like Figure 3A As further shown in the example of FIG4 , the SCQ 46 includes a spatial component scaling unit 252 and a vector quantizer 254. The spatial component scaling unit 252 may receive the spatial component 253 from the ATF encoder 224. The spatial component scaling unit 252 may determine a scaling factor based on the bit allocation 251. For example, the spatial component scaling unit 252 may determine the scaling factor according to the following equation:
[0112]
[0113] In the above equation, the scaling factor (ai) represents the scaling factor of the i-th spatial component 253, where B TOT Denotes the total bit allocation, which is the sum of the bit allocation for the coarse energy and the fine energy corresponding to the i-th transmission channel 225. m,I represents the bit allocation of the i-th instance of the stereo encoder 250.
[0114] For illustration purposes, assume that B TOT = 16 bits and the stereo encoder 250A allocates five (5) bits for the coarse energy (where B C represents the coarse gain bit allocation) and four (4) bits for the fine energy allocation (where B F represents fine gain bit allocation), the spatial component scaling unit 252 can be scaled by the scaling factor a i is determined to be approximately 0.56 (which is approximately equal to nine divided by 16 or 9 / 16). Although described above with respect to the above equation, the spatial component scaling unit 252 may determine the scaling factor in other manners, such as a geometric mean, etc.
[0115] The spatial component scaling unit 252 may apply the scaling factor to the corresponding spatial component in the spatial components 253 to obtain a scaled spatial component 255. The spatial component scaling unit 255 may output the scaled spatial component 255 to the vector quantizer 254. The vector quantizer 254 may perform vector quantization with respect to the scaled spatial component 255 to obtain a quantized spatial component 257.
[0116] PADD 226A may also include a bitstream generator 256 that may receive sub-bitstream 233 and quantized spatial component 257. Bitstream generator 256 may represent a unit configured to specify sub-bitstream 233 and quantized spatial component 257 in bitstream 231. Bitstream 231 may represent the example of bitstream 31 discussed above.
[0117] exist Figure 3B In the example of FIG. 2 , PAED 226B is similar to PAED 226A, except that there is a defined number (i.e., Figure 3B In the example of eight) transmission channels 225a-225h, thereby generating a defined number (ie, Figure 3B In the example of FIG. 2 , four (four in the example) stereo encoders 250 a - 250 d are used. When the transmission channel 225 complies with the HTF, the HTF indicates the presence of eight transmission channels, four of which (e.g., transmission channels 225A - 225D) may define a first-order ambisonic audio signal as a background audio signal including a W ambisonic coefficient (e.g., in transmission channel 225A), an X ambisonic coefficient (e.g., in transmission channel 225B), a Y ambisonic coefficient (e.g., in transmission channel 225C), and a Z ambisonic coefficient (e.g., in transmission channel 225D). The remaining four transmission channels 225E - 225H may each specify a foreground audio signal.
[0118] For the background audio signal, the stereo encoders 250A and 250B may not output any bit allocations, given that there is no corresponding spatial component 253 for the background audio signal. The stereo encoders 250C and 250D may output bit allocations 251C and 251D, which are used by the spatial component scaling unit 252 to determine the scaling factor. Each of the ATF encoder 224, the vector quantizer 254, and the bitstream generator 256 is as described above with respect to Figure 3A The example works as described.
[0119] Next reference Figure 3C In the example of FIG. 2 , PAED 226C is similar to PAED 226A, except that PAED 226C includes redundancy reduction units 280A-280L (“redundancy reduction units 280”) and reconfigures stereo encoder 250 to operate using a differential encoding scheme. PAED 226C may select one of transport channels 225A-225N as a reference transport channel. Figure 3C In the example of FIG, transport channel 225A is provided. PAED 226C provides transport channel 225A to each of stereo encoders 250 as a reference transport channel.
[0120] PAED 226C may also provide transport channel 225A to each of redundancy reduction units 280. Redundancy reduction unit 280 may remove any redundant audio information between transport channel 225A and each of the corresponding remaining transport channels 225B-225M. Redundancy reduction unit 280 may output redundancy-reduced transport channels 281B-281M ("redundancy-reduced transport channels 281") to a corresponding one of stereo encoders 250 after reducing redundancy between reference transport channel 225A and each of the remaining transport channels 225B-225M. Stereo encoder 250 may operate as described above to perform differential encoding with respect to reference transport channel 225A with respect to each of the redundancy-reduced transport channels.
[0121] As a result of the redundancy reduction, PAED 226C may provide better compression efficiency at the expense of additional computational cost (in terms of computational resources, since more stereo encoders 250 may be required compared to PAED 226A). Figure 3C Not shown in the example of FIG, PAED 226C may perform some form of analysis to determine the correlation between reference transport channel 225A and the remaining transport channels 225B-226M. When the correlation is above a certain threshold (indicating relatively high correlation and, therefore, relatively high redundancy), PAED 226C may be used to gain additional compression efficiency. When the correlation is below the threshold, PAED 226A may be invoked because there may not be sufficient compression efficiency to warrant the additional computational cost.
[0122] Figures 4A to 4C are illustrated in more detail in Figure 1 and Figure 2 A block diagram of an example of a psychoacoustic audio decoding device is shown in the example of FIG. Figure 4A , PADD 334A may represent an example of PADD 34 and / or PADD 134. PADD 334A may include a bitstream extractor 338, stereo decoders 340A-340N ("stereo decoder 340"), and SCD 54.
[0123] The bitstream extractor 336 may represent a unit configured to parse the sub-bitstreams 233 and the quantized spatial components 257 from the bitstream 231. The bitstream extractor 338 may output each of the sub-bitstreams 233 to a separate instance of the stereo decoder 340. The bitstream extractor 338 may also output the quantized spatial components 257 to the SCD 54.
[0124] Each of the stereo decoders 340 may reconstruct the second transport channel of the pair of transport channels 225′ based on the quantization gain and quantization shape set forth in the sub-bitstream 233. Each of the stereo decoders 340 may then obtain a residual representing the first transport channel of the pair of transport channels 225′ from the sub-bitstream 233. The stereo decoder 340 may add the residual to the second transport channel to obtain the first transport channel (e.g., transport channel 225A′) from the second transport channel (e.g., transport channel 225B′). The stereo decoder 340 may output the transport channel 225′ to the ATF decoder 336 (which may perform operations similar to (or otherwise substantially similar to) the SADD 36 and / or SDDD 136).
[0125] After dequantizing the quantized gain and the quantized shape, the stereo decoder 340 may determine a bit allocation 251. The bit allocation 251 may specify one or more of a coarse energy bit allocation, a fine energy bit allocation, and a shape bit allocation. As an example, the bit allocation 251 may specify a coarse energy bit allocation and a fine energy bit allocation. The stereo decoder 340 may output the bit allocation 251 to the SCD 56.
[0126] like Figure 4A As further shown in the example of FIG, SCD 56 may include a vector dequantizer 342 and a spatial component descaling unit 344. The vector dequantizer 342 may represent a unit configured to perform the same operations as described above with respect to Figure 3A and Figure 3B The vector quantizer 254 described in the example of FIG. 1 is a unit that operates in an inverse manner. Thus, the vector dequantizer 342 can perform vector dequantization with respect to the quantized spatial component 257′ to obtain the scaled spatial component 255′. The vector dequantizer 342 can output the scaled spatial component 255′ to the spatial component descaling unit 344.
[0127] The spatial component scaling unit 344 may represent a unit configured to descale the scaled spatial component 255' in a manner reciprocal to that described above with respect to the spatial component scaling unit 252. Thus, the spatial component descaling unit 344 may determine the scaling factor based on the bit allocation 251 in the manner described above with respect to the spatial component scaling unit 252. However, rather than multiplying the scaled spatial component 255' by the scaling factor, the spatial component scaling unit 252 may divide the scaled spatial component 255' by the scaling factor to obtain the spatial component 253'. The spatial component descaling unit 344 may output the spatial component 253' to the ATF decoder 336.
[0128] The ATF decoder 336 can receive the transport channels 225' and the spatial components 253' and perform spatial audio decoding with respect to the transport channels 225' and the spatial components to obtain the scene-based audio data 221'. The scene-based audio data 221' can represent an example of the scene-based audio data 211' and / or the scene-based audio data 121'.
[0129] In Figure 4B examples, the PADD 334B is similar to the PADD 334A except that there are a defined number (i.e., eight in the example of Figure 4B ) of transport channels 225A' - 225h' resulting in a defined number (i.e., four in the example of Figure 4B ) of AptX stereo decoders 340A - 340D. When the transport channels 225' conform to the HTF, the HTF indicates that there are eight transport channels, four of which (e.g., transport channels 225A' - 225D') can define a first order ambisonic audio signal as a background audio signal including a W ambisonic coefficient (e.g., in transport channel 225A'), an X ambisonic coefficient (e.g., in transport channel 225B'), a Y ambisonic coefficient (e.g., in transport channel 225C'), and a Z ambisonic coefficient (e.g., in transport channel 225D'). The remaining four transport channels 225E' to 225h' can each specify a foreground audio signal.
[0130] For the background audio signal, the stereo decoders 340A and 340B can not output any bit allocation given that there is no corresponding spatial component 253' for the background audio signal. The stereo decoders 340C and 340D can output bit allocations 251C and 251D, which are used by the spatial component de-scaling unit 344 to determine the scaling factors. Each of the ATF decoder 336, the vector dequantizer 342, and the bitstream extractor 338 function as described above with respect to the example of Figure 4A .
[0131] Referring next to Figure 4C, PADD 334C is similar to PADD 334A, except that PADD 334C operates inversely to PAED 226C and, therefore, includes differential decoding with respect to sub-bitstream 233 and reconstruction synthesis (RS) units 380A-380L ("RS units 380"). Stereo encoder 340 decodes sub-bitstream 233A to output reference transmission channel 225A' and redundancy-reduced transmission channels 281B'-281M' ("redundancy-reduced transmission channels 281"). Stereo decoder 340A may output reference transmission channel 225A' to each of RS units 380, while residual stereo decoders 340B-340M output redundancy-reduced transmission channel 281' to corresponding each of RS units 380. Each of RS units 380 operates inversely to redundancy reduction unit 280 to reintroduce redundancy and thereby reconstruct transmission channels 225B'-225M'.
[0132] Figure 5 is shown in more detail in Figures 3A to 3C The encoder 550 is shown as a multi-channel encoder and represents Figures 3A to 3C The example of the stereo encoder 250 shown in the example of FIG (where the stereo encoder 250 may include only two channels, while the encoder 550 has been generalized to support N channels).
[0133] like Figure 5 As shown in the example of , the encoder 550 includes gain / shape analysis units 552A-552N ("gain / shape analysis units 552"), energy quantization units 556A-556N ("energy quantization units 556"), level difference units 558A-558N ("level difference units 558"), transform units 562A-562N ("transform units 562"), and a vector quantizer 564. Each of the gain / shape analysis units 552 may be as described below with respect to the following. Figure 7 、 Figure 9A and / or Figure 9B The gain-shape analysis unit described in operates to perform a gain-shape analysis with respect to each of the transmission channels 551 to obtain gains 553A-553N ("gains 553") and shapes 555A-555N ("shapes 555").
[0134] The energy quantization unit 556 may be as follows with respect to Figure 7 、 Figure 9A and / or Figure 9BThe energy quantizer 550 operates as described to quantize the gain 553 and thereby obtain quantized gains 557A-557N ("quantized gains 557"). The level difference units 558 may each represent a unit configured to compare a pair of gains 553 to determine a difference between the pair of gains 553. In this example, the level difference unit 558 may compare the reference gain 553A with each of the remaining gains 553 to obtain gain differences 559A-559M ("gain differences 559"). The encoder 550 may specify the quantized reference gain 557A and the gain differences 559 in the bitstream.
[0135] Transform unit 562 may perform subband analysis (as discussed in greater detail below) and apply a transform (e.g., KLT, which refers to Karhunen-Loeve transform) to the subbands of shape 555 to output transformed shapes 563A-563N ("transformed shapes 563"). Vector quantizer 564 may perform vector quantization with respect to transformed shapes 563 to obtain residual IDs 565A-565N ("residual IDs 565"), thereby specifying residual IDs 565 in the bitstream.
[0136] Encoder 550 may also determine a combined bit allocation 560 based on the number of bits allocated to quantized gain 557 and gain difference 559. Combined bit allocation 560 may represent one example of bit allocation 251 discussed in more detail above.
[0137] Figure 6 A more detailed diagram Figures 4A to 4C The decoder 634 is shown as a multi-channel decoder and represents Figure 4A and Figure 4B The example of the stereo decoder 340 is shown in the example of FIG (where the stereo decoder 340 may include only two channels, while the decoder 634 has been generalized to support N channels).
[0138] like Figure 6 As shown in the example of , the decoder 634 includes level combination units 636A-636N ("level combination units 636"), a vector quantizer 638, energy dequantization units 640A-640N ("energy dequantization units 640"), inverse transform units 642A-642N ("transform units 642"), and gain / shape synthesis units 646A-646N ("gain / shape synthesis units 552"). The level combination units 636 may each represent a unit configured to combine the quantized reference gain 553A with each of the gain differences 559 to determine the quantized gain 557.
[0139] The energy dequantization unit 640 may be as follows: Figure 8 、 Figure 10A and / or Figure 10B The energy dequantizer of 550 operates as described above to dequantize the quantized gain 557 to obtain the gain 553'. The encoder 550 may specify the quantized reference gain 557A and the gain difference 559 in the ATF audio data.
[0140] Vector dequantizer 638 may perform vector quantization on residual ID 565 to obtain transformed shape 563′. Transform unit 562 may apply an inverse transform (eg, inverse KLT) to transformed shape 563 and perform subband synthesis (as discussed in more detail below) to output shape 555′.
[0141] Each of the gain / shape synthesis units 552 may be as follows with respect to Figure 7 、 Figure 9A and / or Figure 9B The gain-shape analysis unit 646 operates as described in the example discussed above to perform gain-shape synthesis with respect to each of the gain 553' and the shape 555' to obtain the transmission channel 551'. The gain / shape synthesis unit 646 may output the transmission channel 551' to ATF audio data.
[0142] Encoder 550 may also determine a combined bit allocation 560 based on the number of bits allocated to quantized gain 557 and gain difference 559. Combined bit allocation 560 may represent one example of bit allocation 251 discussed in more detail above.
[0143] Figure 7 is a description of various aspects of the technology configured to perform the Figure 2 The audio encoder 1000A may represent an example of a PAED 126, which may be configured to encode audio data for transmission over a personal area network or "PAN" (e.g., ) for transmission. However, the techniques of the present disclosure performed by the audio encoder 1000A can be used in any context where it is desired to compress audio data. In some examples, the audio encoder 1000A can be configured to encode the audio data 17 according to any of the compression algorithms listed above.
[0144] exist Figure 7 In an example of the present invention, the audio encoder 1000A can be configured to encode the audio data 25 using a gain-shape vector quantization encoding process that includes encoding and decoding residual vectors using a compact mapping. In the gain-shape vector quantization encoding process, the audio encoder 1000A is configured to encode both the gain (e.g., energy level) and shape (e.g., residual vector defined by transform coefficients) of a subband of frequency-domain audio data. Each subband of the frequency-domain audio data represents a certain frequency range of a particular frame of the audio data 25.
[0145] The audio data 25 can be sampled at a specific sampling frequency. Example sampling frequencies can include 48kHz or 44.1kHz, although any desired sampling frequency can be used. Each digital sample of the audio data 25 can be defined by a specific input bit depth (e.g., 16 bits or 24 bits). In one example, the audio encoder 1000A can be configured to operate on a single channel of the audio data 21 (e.g., mono audio). In another example, the audio encoder 1000A can be configured to independently encode two or more channels of the audio data 25. For example, the audio data 17 can include a left channel and a right channel for stereo audio. In this example, the audio encoder 1000A can be configured to independently encode the left audio channel and the right audio channel in dual mono mode. In other examples, the audio encoder 1000A can be configured to encode two or more channels of the audio data 25 together (e.g., in joint stereo mode). For example, the audio encoder 1000A can perform certain compression operations by predicting one channel of the audio data 25 with another channel of the audio data 25.
[0146] Regardless of how the channels of the audio data 25 are arranged, the audio encoder 1000A obtains the audio data 25 and sends the audio data 25 to the transform unit 1100. The transform unit 1100 is configured to transform a frame of the audio data 25 from the time domain to the frequency domain to generate frequency domain audio data 1112. A frame of the audio data 25 can be represented by a predetermined number of samples of audio data. In one example, a frame of the audio data 25 can be 1024 samples wide. Different frame widths can be selected based on the frequency transform used and the desired amount of compression. The frequency domain audio data 1112 can be represented as transform coefficients, where the value of each transform coefficient represents the energy of the frequency domain audio data 1112 at a particular frequency.
[0147] In one example, the transform unit 1100 can be configured to transform the audio data 225 into frequency domain audio data 1112 using a modified discrete cosine transform (MDCT). The MDCT is an "overlapped" transform based on the Type IV discrete cosine transform. The MDCT is considered "overlapped" because it operates on data from multiple frames. That is, to perform a transform using the MDCT, the transform unit 1100 can include a fifty percent overlapping window into subsequent frames of audio data. The overlapping nature of the MDCT can be used in data compression techniques, such as audio coding, because it can reduce artifacts from encoding and decoding at frame boundaries. The transform unit 1100 need not be constrained to using the MDCT, but can use other frequency domain transform techniques to transform the audio data 117 into frequency domain audio data 1112.
[0148] The subband filter 1102 divides the frequency-domain audio data 1112 into subbands 1114. Each of the subbands 1114 includes transform coefficients for the frequency-domain audio data 1112 within a specific frequency range. For example, the subband filter 1102 may divide the frequency-domain audio data 1112 into twenty different subbands. In some examples, the subband filter 1102 may be configured to divide the frequency-domain audio data 1112 into subbands 1114 of uniform frequency range. In other examples, the subband filter 1102 may be configured to divide the frequency-domain audio data 1112 into subbands 1114 of non-uniform frequency range.
[0149] For example, the subband filter 1102 can be configured to divide the frequency-domain audio data 1112 into subbands 1114 according to the Bark scale. Typically, subbands on the Bark scale have frequency ranges that are perceptually equidistant. That is, the subbands on the Bark scale are not equal in terms of frequency range, but are equal in terms of human hearing. Typically, lower-frequency subbands will have fewer transform coefficients because lower frequencies are more easily perceived by the human auditory system. Thus, the frequency-domain audio data 1112 in the lower-frequency subbands of subband 1114 is less compressed by the audio encoder 1000A than in the higher-frequency subbands. Similarly, the higher-frequency subbands of subband 1114 may include more transform coefficients because higher frequencies are more difficult for the human auditory system to perceive. Thus, the frequency-domain audio data 1112 in the higher-frequency subbands of subband 1114 is more compressed by the audio encoder 1000A than in the lower-frequency subbands.
[0150] The audio encoder 1000A can be configured to process each of the subbands 1114 using a subband processing unit 1128. That is, the subband processing unit 1128 can be configured to process each subband separately. The subband processing unit 1128 can be configured to perform a gain-shape vector quantization process with extended range coarse-fine quantization according to the techniques of this disclosure.
[0151] The gain-shape analysis unit 1104 may receive subbands 1114 as input. For each of the subbands 1114, the gain-shape analysis unit 1104 may determine an energy level 1116 for each of the subbands 1114. That is, each subband in the subbands 1114 has an associated energy level 1116. The energy level 1116 is a scalar value in decibels (dBs) that represents the amount of energy (also referred to as the gain) in the transform coefficients for a particular one of the subbands 1114. The gain-shape analysis unit 1104 may separate the energy level 1116 for one of the subbands 1114 from the transform coefficients for the subband to produce a residual vector 1118. The residual vector 1118 represents the so-called "shape" of the subband. The shape of a subband may also be referred to as the spectrum of the subband.
[0152] The vector quantizer 1108 may be configured to quantize the residual vector 1118. In one example, the vector quantizer 1108 may quantize the residual vector using a quantization process to generate a residual ID 1124. Instead of quantizing each sample individually (e.g., scalar quantization), the vector quantizer 1108 may be configured to quantize a block of samples included in the residual vector 1118 (e.g., a shape vector). Any vector quantization technique may be used with the extended range coarse-fine energy quantization process.
[0153] In some examples, the audio encoder 1000A can dynamically allocate bits for encoding and decoding the energy levels 1116 and the residual vectors 1118. That is, for each of the subbands 1114, the audio encoder 1000A can determine the number of bits to allocate for energy quantization (e.g., by the energy quantizer 1106) and the number of bits to allocate for vector quantization (e.g., by the vector quantizer 1108). The total number of bits allocated for energy quantization can be referred to as energy allocation bits. These energy allocation bits can then be allocated between the coarse quantization process and the fine quantization process.
[0154] The energy quantizer 1106 may receive the energy levels 1116 of the subbands 1114 and quantize the energy levels 1116 of the subbands 1114 into coarse energies 1120 and fine energies 1122 (which may represent one or more quantized fine residuals). This disclosure will describe the quantization process for one subband, but it should be understood that the energy quantizer 1106 may perform energy quantization on one or more subbands 1114 (including each of the subbands 1114).
[0155] Typically, the energy quantizer 1106 can perform a recursive two-step quantization process. The energy quantizer 1106 can first quantize the energy level 1116 using a first number of bits for a coarse quantization process to generate a coarse energy 1120. The energy quantizer 1106 can generate the coarse energy using energy levels within a predetermined range for quantization (e.g., a range defined by a maximum and minimum energy level). The coarse energy 1120 approximates the value of the energy level 1116.
[0156] Energy quantizer 1106 can then determine the difference between coarse energy 1120 and energy level 1116. This difference is sometimes referred to as a quantization error. Energy quantizer 1106 can then quantize the quantization error using a second number of bits in a fine quantization process to generate fine energy 1122. The number of bits used for fine quantization is determined by subtracting the number of bits used for the coarse quantization process from the total number of energy allocation bits. When added together, coarse energy 1120 and fine energy 1122 represent the total quantized value of energy level 1116. Energy quantizer 1106 can continue in this manner to generate one or more fine energies 1122.
[0157] The audio encoder 1000A can be further configured to encode the coarse energies 1120, the fine energies 1122, and the residual IDs 1124 using a bitstream encoder 1110 to produce encoded audio data 31 (another way of referring to the bitstream 31). The bitstream encoder 1110 can be configured to use one or more entropy encoding processes to further compress the coarse energies 1120, the fine energies 1122, and the residual IDs 1124. The entropy encoding processes can include Huffman coding, arithmetic coding, context adaptive binary arithmetic coding (CABAC), and other similar encoding techniques.
[0158] In one example of the disclosure, the quantization performed by the energy quantizer 1106 is uniform quantization. That is, the step size (also referred to as the “resolution”) of each quantization is equal. In some examples, the step size can be in decibels (dB). The step size for coarse quantization and fine quantization can be determined by the predetermined range of energy values used for quantization and the number of bits allocated for quantization, respectively. In one example, the energy quantizer 1106 performs uniform quantization for both coarse quantization (e.g., to produce the coarse energies 1120) and fine quantization (e.g., to produce the fine energies 1122).
[0159] Performing a two-step uniform quantization process is equivalent to performing a single uniform quantization process. However, by splitting the uniform quantization into two parts, the bits allocated to coarse quantization and fine quantization can be controlled independently. This can allow for greater flexibility in the allocation of bits across energy and vector quantization, and can improve compression efficiency. Consider an M-level uniform quantizer, where M defines the number of levels (e.g., in dB) that energy levels can be divided into. M can be determined by the number of bits allocated for quantization. For example, the energy quantizer 1106 can use Ml levels for coarse quantization, and M2 levels for fine quantization. This is equivalent to a single uniform quantizer using Ml*M2 levels.
[0160] Figure 8 is a block diagram illustrating a more detailed implementation of a psychoacoustic audio decoder in accordance with the techniques of the present disclosure. The audio decoder 1002A can represent one example of the decoder 510, which can be configured to decode audio data received through a PAN (e.g., the PAN 520). However, the techniques of the present disclosure performed by the audio decoder 1002A can be used in any context in which compressed audio data is desired. In some examples, the audio decoder 1002A can be configured to decode the audio data 21 according to any of the compression algorithms listed above. As such, the techniques of the present disclosure can be used in any audio codec configured to perform quantization of audio data. The audio decoder 1002A can be configured to perform various aspects of the quantization process using compact mapping in accordance with the techniques of the present disclosure. Figures 1 to 3C is a block diagram illustrating a more detailed implementation of a psychoacoustic audio decoder in accordance with the techniques of the present disclosure. The audio decoder 1002A can represent one example of the decoder 510, which can be configured to decode audio data received through a PAN (e.g., the PAN 520). However, the techniques of the present disclosure performed by the audio decoder 1002A can be used in any context in which compressed audio data is desired. In some examples, the audio decoder 1002A can be configured to decode the audio data 21 according to any of the compression algorithms listed above. As such, the techniques of the present disclosure can be used in any audio codec configured to perform quantization of audio data. The audio decoder 1002A can be configured to perform various aspects of the quantization process using compact mapping in accordance with the techniques of the present disclosure.
[0161] Generally, the audio decoder 1002A can operate in a reciprocal manner relative to the audio encoder 1000A. Thus, the same process used in the encoder for quality / bitrate scalable cooperative PVQ can be used in the audio decoder 1002A. Decoding is based on the same principles, the reverse of the operations performed in the decoder, so that the audio data can be reconstructed from the encoded bitstream received from the encoder. Each quantizer has an associated dequantizer counterpart. For example, Figure 8 As shown, the inverse transform unit 1100', the inverse subband filter 1102', the gain-shape synthesis unit 1104', the energy dequantizer 1106', the vector dequantizer 1108' and the bitstream decoder 1110' can be used to perform Figure 7 The inverse operations of the transform unit 1100, subband filter 1102, gain-shape analysis unit 1104, energy quantizer 1106, vector quantizer 1108 and bitstream encoder 1110.
[0162] Specifically, the gain-shape synthesis unit 1104' reconstructs the frequency domain audio data with the reconstructed residual vector and the reconstructed energy level. The inverse subband filter 1102' and the inverse transform unit 1100' output the reconstructed audio data 25'. In examples where the encoding is lossless, the reconstructed audio data 25' may completely match the audio data 25. In examples where the encoding is lossy, the reconstructed audio data 25' may not completely match the audio data 25.
[0163] Figure 9A and Figure 9B is shown in more detail in Figures 1 to 3C A block diagram of another example of a psychoacoustic audio encoder is shown in the example of . Figure 9A For example, the audio encoder 1000B may be configured to encode audio data for transmission via a PAN (e.g., ) is transmitted. However, similarly, the techniques of the present disclosure performed by audio encoder 1000B can be used in any scenario where compressed audio data is desired. In some examples, audio encoder 1000B can be configured to encode audio data 25 according to any of the compression algorithms listed above, including AptX. Thus, the techniques of the present disclosure can be used in any audio codec. As will be explained in more detail below, audio encoder 1000B can be configured to perform various aspects of perceptual audio coding and decoding according to various aspects of the techniques described in this disclosure.
[0164] exist Figure 9AIn an example of the present invention, the audio encoder 1000B can be configured to encode the audio data 25 using a gain-shape vector quantization encoding process. In the gain-shape vector quantization encoding process, the audio encoder 1000B is configured to encode both the gain (e.g., energy level) and shape (e.g., residual vector defined by transform coefficients) of a subband of frequency-domain audio data. Each subband of the frequency-domain audio data represents a certain frequency range of a particular frame of the audio data 25. Generally, throughout this disclosure, the term "subband" refers to a frequency range, a frequency band, or the like.
[0165] The audio encoder 1000B invokes a transform unit 1100 to process the audio data 25. The transform unit 1100 is configured to process the audio data 25 by at least partially applying a transform to frames of the audio data 25 and thereby transforming the audio data 25 from the time domain to the frequency domain to generate frequency domain audio data 1112.
[0166] A frame of audio data 25 may be represented by a predetermined number of samples of audio data. In one example, a frame of audio data 25 may be 1024 samples wide. Different frame widths may be selected based on the frequency transform used and the desired amount of compression. Frequency domain audio data 1112 may be represented as transform coefficients, where the value of each transform coefficient represents the energy of the frequency domain audio data 1112 at a particular frequency.
[0167] In one example, the transform unit 1100 can be configured to transform the audio data 225 into frequency domain audio data 1112 using a modified discrete cosine transform (MDCT). The MDCT is an "overlapped" transform based on the Type IV discrete cosine transform. The MDCT is considered "overlapped" because it operates on data from multiple frames. That is, to perform a transform using the MDCT, the transform unit 1100 may include a fifty percent overlapping window into subsequent frames of audio data. The overlapping nature of the MDCT can be used in data compression techniques, such as audio coding, because it can reduce artifacts from encoding and decoding at frame boundaries. The transform unit 1100 need not be constrained to using the MDCT, but can use other frequency domain transform techniques to transform the audio data 25 into frequency domain audio data 1112.
[0168] The subband filter 1102 divides the frequency-domain audio data 1112 into subbands 1114. Each of the subbands 1114 includes transform coefficients for the frequency-domain audio data 1112 within a specific frequency range. For example, the subband filter 1102 may divide the frequency-domain audio data 1112 into twenty different subbands. In some examples, the subband filter 1102 may be configured to divide the frequency-domain audio data 1112 into subbands 1114 of uniform frequency range. In other examples, the subband filter 1102 may be configured to divide the frequency-domain audio data 1112 into subbands 1114 of non-uniform frequency range.
[0169] For example, the subband filter 1102 can be configured to separate the frequency-domain audio data 1112 into subbands 1114 according to the Bark scale. Typically, subbands on the Bark scale have frequency ranges that are perceptually equidistant. That is, the subbands on the Bark scale are not equal in terms of frequency range, but are equal in terms of human hearing. Typically, lower-frequency subbands will have fewer transform coefficients because lower frequencies are more easily perceived by the human auditory system.
[0170] Thus, the frequency domain audio data 1112 in the lower frequency sub-bands of the sub-band 1114 is compressed less by the audio encoder 1000B than the higher frequency sub-bands. Similarly, the higher frequency sub-bands of the sub-band 1114 may include more transform coefficients because higher frequencies are more difficult for the human auditory system to perceive. Therefore, the frequency domain audio 1112 in the data in the higher frequency sub-bands of the sub-band 1114 may be compressed more by the audio encoder 1000B than the lower frequency sub-bands.
[0171] The audio encoder 1000B may be configured to process each of the subbands 1114 using a subband processing unit 1128. That is, the subband processing unit 1128 may be configured to process each subband individually. The subband processing unit 1128 may be configured to perform a gain-shape vector quantization process.
[0172] Gain-shape analysis unit 1104 may receive subbands 1114 as input. For each of subbands 1114, gain-shape analysis unit 1104 may determine an energy level 1116 for each of subbands 1114. That is, each subband in subbands 1114 has an associated energy level 1116. Energy level 1116 is a scalar value in decibels (dB) that represents the amount of energy (also known as gain) in the transform coefficients for a particular one of subbands 1114. Gain-shape analysis unit 1104 may separate energy level 1116 from the transform coefficients for the subband to produce a residual vector 1118. Residual vector 1118 represents the so-called "shape" of the subband. The shape of the subband may also be referred to as the spectrum of the subband. Vector quantizer 1108 may be configured to quantize residual vector 1118. In one example, vector quantizer 1108 may quantize the residual vector using a quantization process to produce a residual ID 1124. Instead of quantizing each sample individually (eg, scalar quantization), the vector quantizer 1108 may be configured to quantize a block of samples included in the residual vector 1118 (eg, a shape vector).
[0173] In some examples, the audio encoder 1000B can dynamically allocate bits for encoding and decoding the energy levels 1116 and the residual vectors 1118. That is, for each of the subbands 1114, the audio encoder 1000B can determine the number of bits allocated for energy quantization (e.g., by the energy quantizer 1106) and the number of bits allocated for vector quantization (e.g., by the vector quantizer 1108). The total number of bits allocated for energy quantization can be referred to as energy allocation bits. These energy allocation bits can then be allocated between the coarse quantization process and the fine quantization process.
[0174] The energy quantizer 1106 may receive the energy level 1116 of the subband 1114 and quantize the energy level 1116 of the subband 1114 into a coarse energy 1120 and a fine energy 1122. This disclosure will describe the quantization process for one subband, but it should be understood that the energy quantizer 1106 may perform energy quantization on one or more subbands 1114, including each of the subbands 1114.
[0175] like Figure 9A As shown in the example of , the energy quantizer 1106 may include a prediction / difference ("P / D") unit 1130, a coarse quantization ("CQ") unit 1132, a summation unit 1134, and a fine quantization ("FQ") unit 1136. The P / D unit 1130 may predict or otherwise identify the difference between the energy levels 1116 of one of the subbands 1114 and another of the subbands 1114 for the same frame of audio data (which may be referred to as a spatial prediction in the frequency domain) or the same (or possibly different) one of the subbands 1114 from different frames (which may be referred to as a temporal prediction). The P / D unit 1130 may analyze the energy levels 1116 in this manner to obtain a predicted energy level 1131 ("PEL 1131") for each subband 1114. The P / D unit 1130 may output the predicted energy levels 1131 to the coarse quantization unit 1132.
[0176] The coarse quantization unit 1132 may represent a unit configured to perform coarse quantization with respect to the predicted energy level 1131 to obtain the coarse energy 1120. The coarse quantization unit 1132 may output the coarse energy 1120 to the bitstream encoder 1110 and the summing unit 1134. The summing unit 1134 may represent a unit configured to obtain the difference between the coarse quantization unit 1134 and the predicted energy level 1131. The summing unit 1134 may output the difference as an error 1135 (which may also be referred to as a “residual 1135”) to the fine quantization unit 1135.
[0177] The fine quantization unit 1132 may represent a unit configured to perform fine quantization with respect to the error 1135. The fine quantization may be considered "fine" relative to the coarse quantization performed by the coarse quantization unit 1132. That is, the fine quantization unit 1132 may perform quantization according to a step size having a higher resolution than the step size used when performing the coarse quantization, thereby further quantizing the error 1135. As a result of performing the fine quantization with respect to the error 1135, the fine quantization unit 1136 may obtain a fine energy 1122 for each subband 1122. The fine quantization unit 1136 may output the fine energy 1122 to the bitstream encoder 1110.
[0178] Typically, the energy quantizer 1106 can perform a multi-step quantization process. The energy quantizer 1106 can first quantize the energy level 1116 using a first number of bits for a coarse quantization process to generate a coarse energy 1120. The energy quantizer 1106 can generate the coarse energy using energy levels within a predetermined range for quantization (e.g., a range defined by a maximum and minimum energy level). The coarse energy 1120 approximates the value of the energy level 1116.
[0179] The energy quantizer 1106 can then determine the difference between the coarse energy 1120 and the energy level 1116. This difference is sometimes referred to as a quantization error (or residual). The energy quantizer 1106 can then quantize the quantization error using a second number of bits in a fine quantization process to produce a fine energy 1122. The number of bits used for the fine quantization bits is determined by subtracting the number of bits used for the coarse quantization process from the total number of energy allocation bits. When added together, the coarse energy 1120 and the fine energy 1122 represent the total quantized value of the energy level 1116.
[0180] The audio encoder 1000B may be further configured to encode the coarse energy 1120, the fine energy 1122, and the residual ID 1124 using the bitstream encoder 1110 to generate the encoded audio data 21. The bitstream encoder 1110 may be configured to further compress the coarse energy 1120, the fine energy 1122, and the residual ID 1124 using one or more of the entropy encoding processes described above.
[0181] According to aspects of the present disclosure, the energy quantizer 1106 (and / or its components, such as the fine quantization unit 1136) can implement a layered rate control mechanism to provide a greater degree of scalability and enable seamless or substantially seamless real-time streaming. For example, the fine quantization unit 1136 can implement a layered fine quantization scheme according to aspects of the present disclosure. In some examples, the fine quantization unit 1136 invokes a multiplexer (or "MUX") 1137 to implement a selection operation of the layered rate control.
[0182] The term“coarse quantization” refers to the combination operation of the two-step coarse-fine quantization process described above. According to various aspects of the present disclosure, the fine quantization unit 1136 can perform one or more additional iterations of fine quantization with respect to the error 1135 received from the summation unit 1134. The fine quantization unit 1136 can use a multiplexer 1137 to switch between and iterate through various (more) fine levels of energy. The term“fine quantization” refers to the second step of the two-step coarse-fine quantization process described above. According to various aspects of the present disclosure, the fine quantization unit 1136 can perform one or more iterations of fine quantization with respect to the error 1135 received from the summation unit 1134. The fine quantization unit 1136 can use a multiplexer 1137 to switch between and iterate through various (more) fine levels of energy.
[0183] Hierarchical rate control can refer to a tree-based fine quantization structure or a cascaded fine quantization structure. When viewed as a tree-based structure, the existing two-step quantization operation forms a root node of the tree, and the root node is described as having a resolution depth of one (1). Depending on the availability of bits for further fine quantization according to the techniques of the present disclosure, the multiplexer 1137 can select additional levels of fine-grained quantization. With respect to the tree-based structure representing the multi-level fine quantization techniques of the present disclosure, any such subsequent fine quantization level selected by the multiplexer 1137 represents a resolution depth of two (2), three (3), and so on.
[0184] The fine quantization unit 1136 can provide improved scalability and control with respect to seamless real-time streaming scenarios in wireless PANs. For example, the fine quantization unit 1136 can replicate a hierarchical fine quantization scheme and quantization multiplexing tree at a higher level of hierarchy, seeding at the coarse quantization points of a more general decision tree. Moreover, the fine quantization unit 1136 can enable seamless or substantially seamless real-time compression and streaming navigation by the audio encoder 1000B. For example, the fine quantization unit 1136 can perform a multi-root hierarchical decision structure with respect to the multi-level fine quantization, enabling the energy quantizer 1106 to utilize the total available bits to implement potentially several iterations of fine quantization.
[0185] The fine quantization unit 1136 can implement the hierarchical rate control process in various ways. The fine quantization unit 1136 can invoke the multiplexer 1137 on a per-subband basis to independently multiplex (and thereby select a respective tree-based quantization scheme) information relating to each of the subbands 1114 with respect to the error 1135. That is, in these examples, the fine quantization unit 1136 performs a multiplexed hierarchical quantization mechanism selection for each respective subband 1114 independently of the quantization mechanism selection for any other of the subbands 1114. In these examples, the fine quantization unit 1136 quantizes each subband 1114 according to a target bit rate specified with respect to only the respective subband 1114. In these examples, the audio encoder 1000B can signal details of the particular hierarchical quantization scheme for each of the subbands 1114 as part of the encoded audio data 21.
[0186] In other examples, the fine quantization unit 1136 may call the multiplexer 1137 only once and thereby select a single multiplexing-based quantization scheme for the error 1135 information associated with all subbands 1114. That is, in these examples, the fine quantization unit 1136 quantizes the error 1135 information associated with all subbands 1114 according to the same target bit rate, which is selected once and defined uniformly for all subbands 1114. In these examples, the audio encoder 1000B may signal the details of the single hierarchical quantization scheme applied across all subbands 1114 as part of the encoded audio data 21.
[0187] Next reference Figure 9B For example, the audio encoder 1000C may be represented as Figure 1 and Figure 2 Another example of a psychoacoustic audio encoding device 26 and / or 126 is shown in the example of FIG. The audio encoder 1000C is similar to the Figure 9A The audio encoder 1000B shown in the example of , except that the audio encoder 1000C includes a general analysis unit 1148 that can perform gain synthesis analysis or any other type of analysis to output levels 1149 and residuals 1151, a quantization controller unit 1150, a general quantizer 1156 and a cognitive / perceptual / auditory / psychoacoustic (CPHP) quantizer 1160.
[0188] A general analysis unit 1148 may receive the subbands 1114 and perform any type of analysis to generate levels 1149 and residuals 1151 . The general analysis unit 1148 may output the levels 1149 to a quantization controller unit 1150 and the residuals 1151 to a CPHP quantizer 1160 .
[0189] Quantization controller unit 1150 may receive level 1149. Figure 9B As shown in the example of , the quantization controller unit 1150 may include a layer specification unit 1152 and a specification control (SC) manager unit 1154. In response to receiving the level 1149, the quantization controller unit 1150 may call the layer specification unit 1152, which may perform top / bottom / upper layer specification. Figure 11 is a diagram illustrating an example of top-down quantization. Figure 12 , which illustrates an example of bottom-up quantization. That is, the hierarchical specification unit 1152 can switch back and forth between coarse quantization and fine quantization on a frame-by-frame basis to implement a requantization mechanism that can make any given quantization coarser or finer.
[0190] The transition from the coarse state to the finer state may occur by requantizing the previous quantization error. Alternatively, quantization may occur such that adjacent quantization points are grouped together into a single quantization point (moving from the fine state to the coarse state). Such an implementation may use a sequential data structure, such as a linked list or a richer structure, such as a tree or a graph. Thus, the hierarchical specification unit 1152 may determine whether to switch from fine quantization to coarse quantization or from coarse quantization to fine quantization, thereby providing the hierarchical space 1153 (which is the set of quantization points for the current frame) to the SC manager unit 1154. The hierarchical specification unit 1152 may determine whether to switch between finer quantization or coarser quantization based on any information used to perform the fine or coarse quantization specified above (e.g., temporal or spatial priority information).
[0191] The SC manager unit 1154 may receive the layered space 1153 and generate designated metadata 1155, thereby passing an indication 1159 of the layered space 1153 along with the designated metadata 1155 to the bitstream encoder 1110. The SC manager unit 1154 may also output the layered designation 1159 to the quantizer 1156, which may perform quantization with respect to the levels 1149 according to the layered space 1159 to obtain quantized levels 1157. The quantizer 1156 may output the quantized levels 1157 to the bitstream encoder 1110, which may operate as described above to form the encoded audio data 31.
[0192] The CPHP quantizer 1160 may perform one or more of cognitive, perceptual, auditory, and psychoacoustic coding on the residual 1151 to obtain a residual ID 1161. The CPHP quantizer 1160 may output the residual ID 1161 to the bitstream encoder 1110, which may operate as described above to form the encoded audio data 31.
[0193] Figure 10A and Figure 10B A more detailed diagram Figures 1 to 3C A block diagram of another example of a psychoacoustic audio decoder is shown. Figure 10A In the example of Figure 3A1002B. The audio decoder 1002B includes an extraction unit 1232, a subband reconstruction unit 1234, and a reconstruction unit 1236. The extraction unit 1232 may represent a unit configured to extract the coarse energy 1120, the fine energy 1122, and the residual ID 1124 from the encoded audio data 31. The extraction unit 1232 may extract one or more of the coarse energy 1120, the fine energy 1122, and the residual ID 1124 based on the energy bit allocation 1203. The extraction unit 1232 may output the coarse energy 1120, the fine energy 1122, and the residual ID 1124 to the subband reconstruction unit 1234.
[0194] The subband reconstruction unit 1234 may represent a unit configured to operate in a manner reciprocal to the operation of the subband processing unit 1128 of the audio encoder 1000B shown in the example of FIG9 . In other words, the subband reconstruction unit 1234 may reconstruct the subband based on the coarse energy 1120, the fine energy 1122, and the residual ID 1124. The subband reconstruction unit 1234 may include an energy dequantizer 1238, a vector dequantizer 1240, and a subband synthesizer 1242.
[0195] The energy dequantizer 1238 may be configured to Figure 9A 1106. The energy dequantizer 1238 may perform dequantization with respect to the coarse energy 1122 and the fine energy 1122 to obtain a predicted / difference energy level. The energy dequantizer 1238 may perform an inverse prediction or difference calculation to obtain the energy level 1116. The energy dequantizer 1238 may output the energy level 1116 to the subband synthesizer 1242.
[0196] If the encoded audio data 31 includes a syntax element set to a value indicating that the fine energies 1122 are hierarchically quantized, the energy dequantizer 1238 may hierarchically dequantize the fine energies 1122. In some examples, the encoded audio data 31 may include a syntax element indicating whether the hierarchically quantized fine energies 1122 are formed using the same hierarchical quantization structure across all subbands 1114, or whether the respective hierarchical quantization structure is determined separately for each of the subbands 1114. Based on the value of the syntax element, the energy dequantizer 1238 may apply the same hierarchical dequantization structure, as indicated by the fine energies 1122, across all subbands 1114, or may update the hierarchical dequantization structure on a per-subband basis when dequantizing the fine energies 1122.
[0197] The vector dequantizer 1240 may represent a unit configured to perform vector dequantization in a manner reciprocal to the vector quantization performed by the vector quantizer 1108. The vector dequantizer 1240 may perform vector dequantization with respect to the residual ID 1124 to obtain the residual vector 1118. The vector dequantizer 1240 may output the residual vector 1118 to the subband synthesizer 1242.
[0198] The subband synthesizer 1242 may represent a unit configured to operate in an inverse manner to the gain-shape analysis unit 1104. Thus, the subband synthesizer 1242 may perform an inverse gain-shape analysis with respect to the energy levels 1116 and the residual vectors 1118 to obtain the subbands 1114. The subband synthesizer 1242 may output the subbands 1114 to the reconstruction unit 1236.
[0199] The reconstruction unit 1236 may represent a unit configured to reconstruct the audio data 25′ based on the subbands 1114. In other words, the reconstruction unit 1236 may perform inverse subband filtering in a manner reciprocal to the subband filtering applied by the subband filters 1102 to obtain the frequency-domain audio data 1112. The reconstruction unit 1236 may then perform an inverse transform in a manner reciprocal to the transform applied by the transform unit 1100 to obtain the audio data 25′.
[0200] Next reference Figure 10B For example, the audio decoder 1002C may be represented as Figure 1 and / or Figure 2 134. In addition, the audio decoder 1002C can be similar to the audio decoder 1002B, except that the audio decoder 1002C can include an abstract control manager 1250, a hierarchical abstraction unit 1252, a dequantizer 1254, and a CPHP dequantizer 1256.
[0201] The abstract control manager 1250 and the layer abstraction unit 1252 may form a dequantizer controller 1249 that controls the operation of a dequantizer 1254 that operates inversely with the quantizer controller 1150. Thus, the abstract control manager 1250 may operate inversely with the SC manager unit 1154, receiving metadata 1155 and layer designations 1159. The abstract control manager 1250 processes the metadata 1155 and layer designations 1159 to obtain a layer space 1153, which it outputs to the layer abstraction unit 1252. The layer abstraction unit 1252 may operate inversely with the layer designation unit 1152, processing the layer space 1153 to output an indication 1159 of the layer space 1153 to the dequantizer 1254.
[0202] The dequantizer 1254 may operate reciprocally to the quantizer 1156, where the dequantizer 1254 may dequantize the quantized levels 1157 using the indication 1159 of the layered space 1153 to obtain dequantized levels 1149. The dequantizer 1254 may output the dequantized levels 1149 to the subband synthesizer 1242.
[0203] The extraction unit 1232 may output the residual ID 1161 to a CPHP dequantizer 1256, which may operate inversely to the CPHP quantizer 1160. The CPHP dequantizer 1256 may process the residual ID 1161 to dequantize the residual ID 1161 and obtain the residual 1161. The CPHP dequantizer 1256 may output the residual to the subband synthesizer 1242, which may process the residual 1151 and the dequantization level 1254 to output the subband 1114. The reconstruction unit 1236 may operate as described above to convert the subband 1114 into audio data 25′ by applying an inverse subband filter with respect to the subband 1114 and then applying an inverse transform to the output of the inverse subband filter.
[0204] Figure 13 It is shown in the figure Figure 2 A block diagram of example components of a source device is shown in the example of FIG. Figure 13 In the example of FIG1 , source device 112 includes a processor 412, a graphics processing unit (GPU) 414, a system memory 416, a display processor 418, one or more integrated speakers 140, a display 103, a user interface 420, an antenna 421, and a transceiver module 422. In examples where source device 112 is a mobile device, display processor 418 is a mobile display processor (MDP). In some examples (such as examples where source device 112 is a mobile device), processor 412, GPU 414, and display processor 418 can be formed as an integrated circuit (IC).
[0205] For example, an IC can be considered a processing chip within a chip package and can be a system on a chip (SoC). In some examples, two of processor 412, GPU 414, and display processor 418 can be housed together in the same IC, while the other can be housed in a different integrated circuit (i.e., a different chip package), or all three can be housed in different ICs or on the same IC. However, in examples where source device 12 is a mobile device, processor 412, GPU 414, and display processor 418 can all be housed in different integrated circuits.
[0206] Examples of processor 412, GPU 414, and display processor 418 include, but are not limited to, one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Processor 412 may be a central processing unit (CPU) of source device 12. In some examples, GPU 414 may be dedicated hardware that includes integrated and / or discrete logic circuits that provide GPU 414 with massively parallel processing capabilities suitable for graphics processing. In some cases, GPU 414 may also include general-purpose processing capabilities and may be referred to as a general-purpose GPU (GPGPU) when implementing general-purpose processing tasks (i.e., non-graphics-related tasks). Display processor 418 may also be application-specific integrated circuit hardware that is designed to retrieve image content from system memory 416, assemble the image content into image frames, and output the image frames to display 103.
[0207] The processor 412 can execute various types of applications 20. Examples of applications 20 include web browsers, email applications, spreadsheets, video games, other applications that generate visual objects for display, or any of the application types listed in more detail above. The system memory 416 can store instructions for executing the applications 20. Executing one of the applications 20 on the processor 412 causes the processor 412 to generate graphics data for image content to be displayed and audio data 21 to be played (possibly via the integrated speakers 105). The processor 412 can send the graphics data for the image content to the GPU 414 for further processing based on instructions or commands sent by the processor 412 to the GPU 414.
[0208] Processor 412 may communicate with GPU 414 according to a specific application processing interface (API). Examples of such APIs include (Microsoft) API, Khronos Group or OpenGL and OpenCL TM However, aspects of the present disclosure are not limited to DirectX, OpenGL, or OpenCL APIs and may be extended to other types of APIs. Furthermore, the techniques described in this disclosure do not need to work according to an API, and processor 412 and GPU 414 may communicate using any technology.
[0209] System memory 416 can be a memory for source device 12. System memory 416 can include one or more computer-readable storage media. Examples of system memory 416 include, but are not limited to, a random access memory (RAM), an electrically erasable programmable read only memory (EEPROM), a flash memory, or other memory technologies coming after them in the future that are capable of storing and carrying out instructions and / or data structures in the form of desired program code and that are accessible by a computer or processor.
[0210] In some examples, system memory 416 can include instructions that cause processor 412, GPU 414, and / or display processor 418 to perform the functions attributed to processor 412, GPU 414, and / or display processor 418 in this disclosure. Thus, system memory 416 can be a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors (e.g., processor 412, GPU 414, and / or display processor 418) to perform various functions.
[0211] System memory 416 can include a non-transitory storage medium. The term “non-transitory” indicates that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term “non-transitory” should not be interpreted to mean that system memory 416 is non-removable or that its contents are static. As one example, system memory 416 can be removed from source device 12 and moved to another device. As another example, a memory substantially similar to system memory 416 can be inserted into source device 12. In certain examples, a non-transitory storage medium can store data that can change over time (e.g., in RAM).
[0212] User interface 420 can represent one or more hardware or virtual (meaning a combination of hardware and software) user interfaces through which a user can interface with source device 12. User interface 420 can include physical buttons, switches, dials, or virtual versions thereof. User interface 420 can also include a physical or virtual keyboard, a touch interface such as a touchscreen, haptic feedback, etc.
[0213] The processor 412 may include one or more hardware units (including so-called "processing cores") configured to perform all or some portion of the operations discussed above with respect to one or more of the mixing unit 120, audio encoder 122, wireless connection manager 128, and wireless communication unit 130. The antenna 421 and transceiver module 422 may represent units configured to establish and maintain a wireless connection between the source device 12 and the sink device 114. The antenna 421 and transceiver module 422 may represent one or more receivers and / or one or more transmitters capable of wireless communication according to one or more wireless communication protocols. In other words, the transceiver module 422 may represent a separate transmitter, a separate receiver, both a separate transmitter and a separate receiver, or a combined transmitter and receiver. The antenna 421 and transceiver 422 may be configured to receive encoded audio data that has been encoded according to the techniques of the present disclosure. Similarly, the antenna 421 and transceiver 422 may be configured to transmit encoded audio data that has been encoded according to the techniques of the present disclosure. The transceiver module 422 may perform all or some portion of the operations of one or more of the wireless connection manager 128 and the wireless communication unit 130 .
[0214] Figure 14 It is shown in the figure Figure 2 Although the terminal device 114 may include the same components as described above with respect to Figure 13 Although sink device 114 may include components similar to those of source device 112 discussed in greater detail above as an example, sink device 114 may, in some cases, include only a subset of the components discussed above with respect to source device 112 .
[0215] exist Figure 14 In the example of FIG. 1 , terminal device 114 includes one or more speakers 802, a processor 812, a system memory 816, a user interface 820, an antenna 821, and a transceiver module 822. Processor 812 may be similar to or substantially similar to processor 812. In some cases, processor 812 may differ from processor 412 in terms of overall processing power or may be customized for low power consumption. System memory 816 may be similar to or substantially similar to system memory 416. Speaker 140, user interface 820, antenna 821, and transceiver module 822 may be similar to or substantially similar to corresponding speakers 440, user interface 420, and transceiver module 422. Terminal device 114 may also optionally include a display 800, although display 800 may represent a low-power, low-resolution (potentially black and white LED) display used to convey limited information, which may be driven directly by processor 812.
[0216] The processor 812 may include one or more hardware units (including so-called "processing cores") configured to perform all or some portion of the operations discussed above with respect to one or more of the wireless connection manager 150, the wireless communication unit 152, and the audio decoder 132. The antenna 821 and the transceiver module 822 may represent units configured to establish and maintain a wireless connection between the source device 112 and the sink device 114. The antenna 821 and the transceiver module 822 may represent one or more receivers and one or more transmitters capable of wireless communication according to one or more wireless communication protocols. The antenna 821 and the transceiver 822 may be configured to receive encoded audio data that has been encoded according to the techniques of the present disclosure. Similarly, the antenna 821 and the transceiver 822 may be configured to transmit encoded audio data that has been encoded according to the techniques of the present disclosure. The transceiver module 822 may perform all or some portion of the operations of one or more of the wireless connection manager 150 and the wireless communication unit 152.
[0217] Figure 15 It is shown in the figure Figure 1 Flowchart of example operation of an audio encoder in performing various aspects of the techniques described in this disclosure, as shown in the example of FIG. 1 . In operation, the audio encoder 22 may invoke a spatial audio encoding device 24, which may perform spatial audio encoding on the scene-based audio data 21 to obtain a foreground audio signal and corresponding spatial components (1300). Thus, the spatial audio encoding performed by the spatial audio encoding device 24 omits the spatial component quantization described above, as the quantization is again offloaded to the psychoacoustic audio encoding device 26. The spatial audio encoding device 24 may output the ATF audio data 25 to the psychoacoustic audio encoding device 26.
[0218] The audio encoder 22 invokes the psychoacoustic audio encoding device 26 to perform psychoacoustic audio encoding on the foreground audio signal to obtain an encoded foreground audio signal (1302). The psychoacoustic audio encoding device 26 may determine a bit allocation for the foreground audio signal when performing psychoacoustic audio encoding on the foreground audio signal (1304). The psychoacoustic audio encoding device 26 may invoke the SCQ 46, thereby passing the bit allocation to the SCQ 46. The SCQ 46 may scale the spatial components based on the bit allocation for the foreground audio signal to obtain scaled spatial components (1306). The SCQ 46 may then quantize (e.g., vector quantize) the scaled spatial components to obtain quantized spatial components (1308). The psychoacoustic audio encoding device 26 may then specify the encoded foreground audio signal and the quantized spatial components in the bitstream 31 (1310).
[0219] Figure 16 It is shown in the figure Figure 11402.
[0220] In any case, when psychoacoustic audio encoding is performed with respect to the foreground audio signal, the psychoacoustic audio decoding device 34 may determine a bit allocation for the encoded foreground audio signal (1404). The psychoacoustic audio decoding device 34 may call the SCD 54, thereby communicating the bit allocation to the SCD 54. The SCD 54 may descale the scaled spatial components based on the bit allocation for the foreground audio signal to obtain quantized spatial components (1406). The SCD 54 may then dequantize (e.g., vector dequantize) the scaled spatial components to obtain spatial components (1408). The psychoacoustic audio decoding device 34 may reconstruct the ATF audio data 25' based on the foreground audio signal and the spatial components. The spatial audio decoding device 36 may then reconstruct the scene-based audio data 21' based on the foreground audio signal and the spatial components of the ATF audio data 25' (1410).
[0221] The foregoing aspects of these techniques may be implemented in accordance with the following clauses.
[0222] Item 1D. A device configured to encode scene-based audio data, the device comprising: a memory configured to store the scene-based audio data; and one or more processors configured to: perform spatial audio encoding relative to the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; perform psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; determine a bit allocation for the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; scale the spatial components based on the bit allocation for the foreground audio signal to obtain scaled spatial components; quantize the scaled spatial components to obtain quantized spatial components; and specify the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0223] Clause 2D. The device of Clause 1D, wherein the one or more processors are configured to perform psychoacoustic audio encoding according to an AptX compression algorithm with respect to the foreground audio signal to obtain an encoded foreground audio signal.
[0224] Clause 3D. A device as described in any combination of clauses 1D to 2D, wherein the one or more processors are configured to: perform shape and gain analysis with respect to a foreground audio signal to obtain a shape and gain representative of the foreground audio signal; perform quantization with respect to the gain to obtain a coarse quantization gain and one or more fine quantization residuals; and scale the spatial component based on a number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain a scaled spatial component.
[0225] Clause 4D. The apparatus of any combination of clauses 1D to 3D, wherein the one or more processors are configured to perform a linear reversible transform on the scene-based audio data to obtain the foreground audio signal and the corresponding spatial component.
[0226] Clause 5D. The apparatus of any combination of clauses 1D to 4D, wherein the scene-based audio data comprises ambisonic reverberation coefficients corresponding to an order greater than one.
[0227] Clause 6D. The apparatus of any combination of clauses 1D to 4D, wherein the scene-based audio data includes ambisonic reverberation coefficients corresponding to orders greater than zero.
[0228] Clause 7D. The apparatus of any combination of clauses 1D through 6D, wherein the scene-based audio data comprises audio data defined in a spherical harmonics domain.
[0229] Clause 8D. The apparatus of any combination of clauses 1D to 7D, wherein the foreground audio signal comprises a foreground audio signal defined in a spherical harmonics domain, and wherein the spatial components comprise spatial components defined in the spherical harmonics domain.
[0230] Clause 9D. The apparatus of any combination of clauses 1D to 8D, wherein the scene-based audio data comprises multi-order ambisonic audio data.
[0231] Item 10D. A method for encoding scene-based audio data, the method comprising: performing spatial audio encoding relative to the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; performing psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; determining a bit allocation for the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; scaling the spatial components based on the bit allocation for the foreground audio signal to obtain scaled spatial components; quantizing the scaled spatial components to obtain quantized spatial components; and specifying the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0232] Clause 11D. The method of Clause 10D, wherein performing psychoacoustic audio encoding comprises performing psychoacoustic audio encoding according to an AptX compression algorithm with respect to the foreground audio signal to obtain an encoded foreground audio signal.
[0233] Clause 12D. A method as described in any combination of clauses 10D to 11D, wherein performing psychoacoustic audio encoding includes: performing shape and gain analysis with respect to a foreground audio signal to obtain a shape and gain representing the foreground audio signal; performing quantization with respect to the gain to obtain a coarse quantization gain and one or more fine quantization residuals; and scaling the spatial component based on a number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain a scaled spatial component.
[0234] Clause 13D. The method of any combination of clauses 10D to 12D, wherein performing spatial audio encoding comprises performing a linear reversible transform with respect to the scene-based audio data to obtain a foreground audio signal and a corresponding spatial component.
[0235] Clause 14D. The method of any combination of clauses 10D to 13D, wherein the scene-based audio data includes ambisonic reverberation coefficients corresponding to an order greater than one.
[0236] Clause 15D. The method of any combination of clauses 10D to 13D, wherein the scene-based audio data includes ambisonic reverberation coefficients corresponding to orders greater than zero.
[0237] Clause 16D. The method of any combination of clauses 10D to 15D, wherein the scene-based audio data comprises audio data defined in a spherical harmonics domain.
[0238] Clause 17D. The method of any combination of clauses 10D to 16D, wherein the foreground audio signal comprises a foreground audio signal defined in a spherical harmonics domain, and wherein the spatial components comprise spatial components defined in the spherical harmonics domain.
[0239] Clause 18D. The method of any combination of clauses 10D to 17D, wherein the scene-based audio data comprises multi-order ambisonic audio data.
[0240] Item 19D. A device configured to encode scene-based audio data, the device comprising: a component for performing spatial audio encoding relative to the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; a component for performing psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; a component for determining a bit allocation to the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; a component for scaling the spatial components based on the bit allocation to the foreground audio signal to obtain scaled spatial components; a component for quantizing the scaled spatial components to obtain quantized spatial components; and a component for specifying the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0241] Clause 20D. The apparatus of Clause 19D, wherein the means for performing psychoacoustic audio encoding comprises means for performing psychoacoustic audio encoding according to an AptX compression algorithm with respect to the foreground audio signal to obtain an encoded foreground audio signal.
[0242] Clause 21D. An apparatus as described in any combination of clauses 19D to 20D, wherein the means for performing psychoacoustic audio encoding includes: means for performing shape and gain analysis relative to a foreground audio signal to obtain a shape and gain representative of the foreground audio signal; and means for performing quantization relative to the gain to obtain a coarse quantization gain and one or more fine quantization residuals, and wherein the means for scaling the spatial component includes means for scaling the spatial component based on a number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain a scaled spatial component.
[0243] Clause 22D. The apparatus of any combination of clauses 19D to 21D, wherein the means for performing spatial audio encoding comprises means for performing a linear reversible transform with respect to the scene-based audio data to obtain a foreground audio signal and a corresponding spatial component.
[0244] Clause 23D. The apparatus of any combination of clauses 19D to 22D, wherein the scene-based audio data comprises ambisonic reverberation coefficients corresponding to an order greater than one.
[0245] Clause 24D. The apparatus of any combination of clauses 19D to 22D, wherein the scene-based audio data includes ambisonic reverberation coefficients corresponding to orders greater than zero.
[0246] Clause 25D. The apparatus of any combination of clauses 19D to 24D, wherein the scene-based audio data comprises audio data defined in a spherical harmonics domain.
[0247] Clause 26D. The apparatus of any combination of clauses 19D to 25D, wherein the foreground audio signal comprises a foreground audio signal defined in a spherical harmonics domain, and wherein the spatial components comprise spatial components defined in the spherical harmonics domain.
[0248] Clause 27D. The apparatus of any combination of clauses 19D to 26D, wherein the scene-based audio data comprises multi-order ambisonic audio data.
[0249] Item 28D. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: perform spatial audio encoding relative to scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; perform psychoacoustic audio encoding relative to the foreground audio signal to obtain an encoded foreground audio signal; determine a bit allocation for the foreground audio signal when performing psychoacoustic audio encoding relative to the foreground audio signal; scale the spatial components based on the bit allocation for the foreground audio signal to obtain scaled spatial components; quantize the scaled spatial components to obtain quantized spatial components; and specify the encoded foreground audio signal and the quantized spatial components in a bitstream.
[0250] Item 1E. A device configured to decode a bitstream representing encoded scene-based audio data, the device comprising: a memory configured to store the bitstream, the bitstream comprising an encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of the encoded foreground audio signal; and one or more processors configured to: perform psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain a foreground audio signal; determine a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; dequantize the quantized spatial components to obtain scaled spatial components; descale the scaled spatial components based on the bit allocation for the encoded foreground audio signal to obtain spatial components; and reconstruct the scene-based audio data based on the foreground audio signal and the spatial components.
[0251] Clause 2E. The device of Clause 1E, wherein the one or more processors are configured to perform psychoacoustic audio decoding according to an AptX compression algorithm with respect to the encoded foreground audio signal to obtain the foreground audio signal.
[0252] Clause 3E. An apparatus as described in any combination of clauses 1E to 2E, wherein the one or more processors are configured to: obtain from a bitstream a number of bits allocated to each of a coarse quantization gain and one or more fine quantization residuals, the coarse quantization gain and the one or more fine quantization residuals representing a gain of a foreground audio signal, and descale the scaled spatial component based on the number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain a spatial component.
[0253] Clause 4E. The apparatus of any combination of clauses 1E to 3E, wherein the scene-based audio data comprises ambisonic reverberation coefficients corresponding to spherical basis functions having an order greater than zero.
[0254] Clause 5E. The apparatus of any combination of clauses 1E to 4E, wherein the scene-based audio data comprises higher-order ambisonic reverberation coefficients corresponding to an order greater than one.
[0255] Clause 6E. The apparatus of any combination of clauses 1E to 4E, wherein the scene-based audio data comprises audio data defined in a spherical harmonics domain.
[0256] Clause 7E. The apparatus of any combination of clauses 1E to 6E, wherein the encoded foreground audio signal comprises an encoded foreground audio signal defined in a spherical harmonics domain, and wherein the scaled spatial components comprise scaled spatial components defined in a spherical harmonics domain.
[0257] Clause 8E. A device as described in any combination of clauses 1E to 7E, wherein the one or more processors are further configured to: render the scene-based audio data to one or more speaker feeds; and reproduce the sound field represented by the scene-based audio data based on the speaker feeds.
[0258] Clause 9E. A device as described in any combination of clauses 1E to 7E, wherein the one or more processors are further configured to render the scene-based audio data to one or more speaker feeds; and wherein the device includes one or more speakers configured to reproduce the sound field represented by the scene-based audio data based on the speaker feeds.
[0259] Clause 10E. The apparatus of any combination of clauses 1E to 9E, wherein the scene-based audio data comprises multi-order ambisonic audio data.
[0260] Item 11E. A method for decoding a bitstream representing scene-based audio data, the method comprising: obtaining an encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of the encoded foreground audio signal from the bitstream; performing psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain the foreground audio signal; determining a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; dequantizing the quantized spatial components to obtain scaled spatial components; descaling the scaled spatial components based on the bit allocation for the encoded foreground audio signal to obtain spatial components; and reconstructing the scene-based audio data based on the foreground audio signal and the spatial components.
[0261] Clause 12E. The method of Clause 11E, wherein performing psychoacoustic audio decoding comprises performing psychoacoustic audio decoding according to an AptX compression algorithm with respect to the encoded foreground audio signal to obtain the foreground audio signal.
[0262] Clause 13E. A method as described in any combination of clauses 11E to 21E, wherein determining the bit allocation includes obtaining from a bitstream a number of bits allocated to each of a coarse quantization gain and one or more fine quantization residuals, the coarse quantization gain and the one or more fine quantization residuals representing a gain of a foreground audio signal, and wherein descaling the scaled spatial component includes descaling the scaled spatial component based on the number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain the spatial component.
[0263] Clause 14E. The method of any combination of clauses 11E to 13E, wherein the scene-based audio data comprises ambisonic reverberation coefficients corresponding to spherical basis functions having an order greater than zero.
[0264] Clause 15E. The method of any combination of clauses 11E to 14E, wherein the scene-based audio data includes higher-order ambisonic reverberation coefficients corresponding to an order greater than one.
[0265] Clause 16E. The method of any combination of clauses 11E to 14E, wherein the scene-based audio data comprises audio data defined in a spherical harmonics domain.
[0266] Clause 17E. The method of any combination of clauses 11E to 16E, wherein the encoded foreground audio signal comprises an encoded foreground audio signal defined in a spherical harmonics domain, and wherein the scaled spatial components comprise scaled spatial components defined in a spherical harmonics domain.
[0267] Clause 18E. The method of any combination of clauses 11E to 17E, wherein the scene-based audio data is rendered to one or more speaker feeds; and the sound field represented by the scene-based audio data is reproduced based on the speaker feeds.
[0268] Clause 19E. The method of any combination of clauses 11E to 19E, wherein the scene-based audio data comprises multi-order ambisonic audio data.
[0269] Item 20E. A device configured to decode a bitstream representing encoded scene-based audio data, the device comprising: means for obtaining an encoded foreground audio signal and corresponding scaled spatial components defining spatial characteristics of the encoded foreground audio signal from the bitstream; means for performing psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain the foreground audio signal; means for determining a bit allocation to the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; means for dequantizing the quantized spatial components to obtain scaled spatial components; means for descaling the scaled spatial components based on the bit allocation to the encoded foreground audio signal to obtain spatial components; and means for reconstructing scene-based audio data based on the foreground audio signal and the spatial components.
[0270] Clause 21E. The apparatus of Clause 20E, wherein the means for performing psychoacoustic audio decoding comprises means for performing psychoacoustic audio decoding according to an AptX compression algorithm with respect to the encoded foreground audio signal to obtain the foreground audio signal.
[0271] Clause 22E. An apparatus as described in any combination of clauses 20E to 21E, wherein the means for determining a bit allocation includes means for obtaining from a bitstream a number of bits allocated to each of a coarse quantization gain and one or more fine quantization residuals, the coarse quantization gain and the one or more fine quantization residuals representing a gain of a foreground audio signal, and wherein the means for descaling the scaled spatial component includes means for descaling the scaled spatial component based on the number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain the spatial component.
[0272] Clause 23E. The apparatus of any combination of clauses 20E to 22E, wherein the scene-based audio data comprises ambisonic reverberation coefficients corresponding to spherical basis functions having an order greater than zero.
[0273] Clause 24E. The apparatus of any combination of clauses 20E to 23E, wherein the scene-based audio data comprises higher-order ambisonic reverberation coefficients corresponding to an order greater than one.
[0274] Clause 25E. The apparatus of any combination of clauses 20E to 23E, wherein the scene-based audio data comprises audio data defined in a spherical harmonics domain.
[0275] Clause 26E. The apparatus of any combination of clauses 20E to 25E, wherein the encoded foreground audio signal comprises an encoded foreground audio signal defined in a spherical harmonics domain, and wherein the scaled spatial components comprise scaled spatial components defined in a spherical harmonics domain.
[0276] Clause 27E. The apparatus of any combination of clauses 20E to 26E, further comprising means for rendering scene-based audio data to one or more speaker feeds; and means for reproducing a sound field represented by the scene-based audio data based on the speaker feeds.
[0277] Clause 28E. The apparatus of any combination of clauses 20E to 28E, wherein the scene-based audio data comprises multi-order ambisonic audio data.
[0278] Item 29E. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: obtain an encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of the encoded foreground audio signal from a bitstream representing scene-based audio data; perform psychoacoustic audio decoding relative to the encoded foreground audio signal to obtain the foreground audio signal; determine a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding relative to the encoded foreground audio signal; dequantize the quantized spatial components to obtain scaled spatial components; descale the scaled spatial components based on the bit allocation to the encoded foreground audio signal to obtain spatial components; and reconstruct the scene-based audio data based on the foreground audio signal and the spatial components.
[0279] In some scenarios, such as broadcast scenarios, the audio encoding device can be divided into a spatial audio encoder that performs a form of intermediate compression with respect to the ambisonic reverberation representation including gain control, and a psychoacoustic audio encoder 26 (which may also be referred to as "perceptual audio encoder 26") that performs perceptual audio compression to reduce redundancy in data between gain normalized transmission channels.
[0280] Furthermore, the aforementioned techniques can be implemented for any number of different scenarios and audio ecosystems and should not be limited to any of the scenarios or audio ecosystems described above. Several example scenarios are described below, although the techniques should not be limited to these example scenarios. An example audio ecosystem may include audio content, a movie studio, a music studio, a game audio studio, channel-based audio content, a codec engine, a game audio backbone, a game audio codec / rendering engine, and a delivery system.
[0281] Movie studios, music studios, and game audio studios can receive audio content. In some examples, the audio content can represent the output of the capture. Movie studios can output channel-based audio content (e.g., in 2.0, 5.1, and 7.1), such as by using a digital audio workstation (DAW). Music studios can output channel-based audio content (e.g., 2.0 and 5.1), such as by using a DAW. In either case, an encoding engine can receive channel-based audio content based on one or more codecs (e.g., AAC, AC3, Dolby True HD, Dolby Digital Plus, and DTS Master Audio) and encode it for output by a delivery system. A game audio studio can output one or more game audio stems, such as by using a DAW. A game audio codec / rendering engine can encode and / or render the audio stems into channel-based audio content for output by a delivery system. Another example scenario in which the described technology can be performed includes an audio ecosystem, which can include broadcast recorded audio objects, professional audio systems, capture on consumer devices, ambisonic audio formats, on-device rendering, consumer audio, TVs and accessories, and car audio systems.
[0282] Broadcast recorded audio objects, professional audio systems, and capture on consumer devices can all encode their output using the Ambisonics audio format. In this way, the audio content can be encoded and decoded using the Ambisonics audio format into a single representation that can be played back using on-device rendering, consumer audio, TV and accessories, and car audio systems. In other words, a single representation of the audio content can be played back on a universal audio playback system such as audio playback system 16 (i.e., as opposed to requiring a specific configuration such as 5.1, 7.1, etc.).
[0283] Other examples of contexts in which the described techniques may be implemented include audio ecosystems, which may include capture elements and playback elements. The capture elements may include wired and / or wireless capture devices (e.g., Eigen microphones), on-device surround sound capture, and mobile devices (e.g., smartphones and tablets). In some examples, the wired and / or wireless capture devices may be coupled to the mobile devices via (one or more) wired and / or wireless communication channels.
[0284] According to one or more techniques of this disclosure, a mobile device can be used to capture a soundfield. For example, a mobile device can capture a soundfield via a wired and / or wireless capture device and / or an on-device surround sound capturer (e.g., a plurality of microphones integrated into the mobile device). The mobile device can then encode-decode the captured soundfield into ambisonic coefficients for playback by one or more playback elements. For example, a user of a mobile device can record a live event (e.g., a meeting, a panel, a performance, a concert, etc.) (capturing a soundfield of the live event) and encode-decode the recording into ambisonic coefficients.
[0285] A mobile device can also utilize one or more playback elements to playback an ambisonic encoded soundfield. For example, a mobile device can decode an ambisonic encoded soundfield and output a signal to one or more playback elements that causes the one or more playback elements to reproduce the soundfield. As one example, a mobile device can utilize a wireless and / or a wireless communication channel to output a signal to one or more speakers (e.g., a speaker array, a soundbar, etc.). As another example, a mobile device can utilize a docking solution to output a signal to one or more docking stations and / or one or more docking speakers (e.g., a sound system in a smart car and / or a home). As another example, a mobile device can utilize a headphone rendering to output a signal to a set of headphones, for example, to produce a realistic binaural sound.
[0286] In some examples, a particular mobile device can capture a 3D soundfield and later playback the same 3D soundfield. In some examples, a mobile device can capture a 3D soundfield, encode the 3D soundfield into HOA, and send the encoded 3D soundfield to one or more other devices (e.g., other mobile devices and / or other non-mobile devices) for playback.
[0287] Yet another context in which the techniques can be performed includes an audio ecosystem, which can include audio content, a game studio, encoded audio content, a rendering engine, and a delivery system. In some examples, the game studio can include one or more DAWs that can support editing of ambisonic signals. For example, the one or more DAWs can include ambisonic plugins and / or tools that can be configured to operate with (e.g., work with) one or more game audio systems. In some examples, the game studio can output a new backbone format that supports HOA. In any case, the game studio can output encoded audio content to a rendering engine, which can render a soundfield for playback by a delivery system.
[0288] The techniques may also be performed for an exemplary audio capture device. For example, the techniques may be performed for an Eigen microphone, which may include multiple microphones collectively configured to record a 3D sound field. In some examples, the multiple microphones of the Eigen microphone may be located on the surface of a substantially spherical ball having a radius of approximately 4 cm. In some examples, the audio encoding device 20 may be integrated into the Eigen microphone to output the bitstream 21 directly from the microphone.
[0289] Another exemplary audio acquisition scenario may include a production truck, which may be configured to receive signals from one or more microphones, such as one or more Eigen microphones. The production truck may also include an audio encoder, such as Figure 1 The spatial audio encoder device 24.
[0290] In some instances, the mobile device may also include multiple microphones that are collectively configured to record a 3D sound field. In other words, the multiple microphones may have X, Y, Z diversity. In some examples, the mobile device may include a microphone that can be rotated to provide X, Y, Z diversity relative to one or more other microphones of the mobile device. The mobile device may also include an audio encoder, such as Figure 1 Audio encoder 22.
[0291] The enhanced video capture device can be further configured to record a 3D sound field. In some examples, the enhanced video capture device can be attached to the helmet of a user participating in an activity. For example, the enhanced video capture device can be attached to the user's helmet while rafting. In this way, the enhanced video capture device can capture a 3D sound field representing the action around the user (e.g., a collision behind the user, another rafter talking in front of the user, etc.).
[0292] The techniques can also be performed for an accessory-enhanced mobile device that can be configured to record a 3D sound field. In some examples, the mobile device can be similar to the mobile device discussed above, with one or more accessories added. For example, an Eigen microphone can be attached to the mobile device described above to form an accessory-enhanced mobile device. In this way, the accessory-enhanced mobile device can capture a higher-quality version of the 3D sound field than using only the sound capture components integrated into the accessory-enhanced mobile device.
[0293] An example audio playback device that can perform various aspects of the techniques described in this disclosure is discussed further below. According to one or more techniques of this disclosure, speakers and / or soundbars can be arranged in any arbitrary configuration while still playing back a 3D sound field. Additionally, in some examples, the headphone playback device can be coupled to a decoder 32 (which refers to a decoder) via a wired or wireless connection. Figure 1According to one or more techniques of this disclosure, a single universal representation of a sound field can be used to render the sound field on any combination of speaker, soundbar, and headphone playback devices.
[0294] Several different example audio playback environments may also be suitable for performing various aspects of the techniques described in this disclosure. For example, a 5.1 speaker playback environment, a 2.0 (e.g., stereo) speaker playback environment, a 9.1 speaker playback environment with full-height front speakers, a 22.2 speaker playback environment, a 16.0 speaker playback environment, a car speaker playback environment, and a mobile device with an earbud playback environment may be suitable environments for performing various aspects of the techniques described in this disclosure.
[0295] According to one or more techniques of this disclosure, a single universal representation of a sound field can be utilized to render the sound field on any of the aforementioned playback environments. Additionally, the techniques of this disclosure enable a renderer to render a sound field from the universal representation for playback on a playback environment that is different from the environment described above. For example, if design considerations prohibit proper placement of speakers according to a 7.1 speaker playback environment (e.g., if placement of a right surround speaker is not possible), then the techniques of this disclosure enable the renderer to compensate for the other six speakers so that playback can be achieved on a 6.1 speaker playback environment.
[0296] In addition, the user can watch a sports game while wearing headphones. According to one or more techniques of the present disclosure, a 3D sound field of a sports game can be collected (for example, one or more Eigen microphones can be placed in and / or around a baseball stadium), a stereo reverberation coefficient corresponding to the 3D sound field can be obtained and transmitted to a decoder, the decoder can reconstruct the 3D sound field based on the stereo reverberation coefficient and output the reconstructed 3D sound field to a renderer, and the renderer can obtain an indication of the type of playback environment (for example, a headphone) and render the reconstructed 3D sound field into a signal that causes the headphone to output a representation of the 3D sound field of the sports game.
[0297] In each of the various examples described above, it should be understood that the audio encoding device 22 may perform a method or otherwise include components for performing each step of the method that the audio encoding device 22 is configured to perform. In some examples, the components may include one or more processors. In some examples, the one or more processors may represent dedicated processors configured by instructions stored on a non-transitory computer-readable storage medium. In other words, various aspects of the technology in each of the multiple sets of coding examples may provide a non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause the one or more processors to perform the method that the audio encoding device 20 has been configured to perform.
[0298] In one or more examples, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to tangible media such as data storage media. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in this disclosure. A computer program product can include a computer-readable medium.
[0299] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but are instead directed to non-temporary tangible storage media. As used herein, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks, where disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0300] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), processing circuits (including fixed-function circuits and / or programmable processing circuits), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software units configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the techniques may be implemented entirely in one or more circuits or logic elements.
[0301] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., a chipset). Various components, units, or elements are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units can be combined in a codec hardware unit in conjunction with appropriate software and / or firmware or provided by a collection of interoperable hardware units, including one or more processors as described above.
[0302] Furthermore, as used herein, "A and / or B" means "A or B", or "both A and B".
[0303] Various aspects of the technology have been described. These and other aspects of the technology are within the scope of the following claims.
Claims
1. A device configured to encode scene-based audio data, the device comprising: a memory configured to store the scene-based audio data; as well as One or more processors configured to: performing spatial audio coding on the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; performing psychoacoustic audio coding with respect to the foreground audio signal to obtain an encoded foreground audio signal; determining a bit allocation for the foreground audio signal when performing psychoacoustic audio coding with respect to the foreground audio signal; scaling the spatial component based on the bit allocation to the foreground audio signal to obtain a scaled spatial component; quantizing the scaled spatial components to obtain quantized spatial components; as well as The encoded foreground audio signal and the quantized spatial components are specified in a bitstream.
2. The device according to claim 1, wherein The one or more processors are configured to perform psychoacoustic audio coding according to a compression algorithm with respect to the foreground audio signal to obtain the encoded foreground audio signal.
3. The apparatus of claim 1, wherein: The one or more processors are configured to: performing a shape and gain analysis with respect to the foreground audio signal to obtain a shape and gain representative of the foreground audio signal; performing quantization with respect to the gain to obtain a coarse quantization gain and one or more fine quantization residuals; as well as The spatial component is scaled based on a number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain the scaled spatial component.
4. The apparatus of claim 1, wherein: The one or more processors are configured to perform a linear reversible transform on the scene-based audio data to obtain the foreground audio signal and the corresponding spatial component.
5. The apparatus of claim 1, wherein: The scene-based audio data includes ambisonic reverberation coefficients corresponding to an order greater than one.
6. The apparatus of claim 1, wherein: The scene-based audio data includes ambisonic reverberation coefficients corresponding to orders greater than zero.
7. The apparatus of claim 1, wherein: The scene-based audio data includes audio data defined in a spherical harmonics domain.
8. The device according to claim 7, in, The foreground audio signal comprises a foreground audio signal defined in the spherical harmonics domain, and The spatial components include spatial components defined in the spherical harmonics domain.
9. The apparatus of claim 1, wherein: The scene-based audio data includes mixed-order ambisonic reverberation audio data.
10. A method for encoding scene-based audio data, the method comprising: performing spatial audio coding on the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; performing psychoacoustic audio coding with respect to the foreground audio signal to obtain an encoded foreground audio signal; determining a bit allocation for the foreground audio signal when performing psychoacoustic audio coding with respect to the foreground audio signal; scaling the spatial component based on the bit allocation to the foreground audio signal to obtain a scaled spatial component; quantizing the scaled spatial components to obtain quantized spatial components; as well as The encoded foreground audio signal and the quantized spatial components are specified in a bitstream.
11. A device configured to decode a bitstream representing encoded scene-based audio data, the device comprising: a memory configured to store the bitstream, the bitstream comprising the encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of the encoded foreground audio signal; as well as One or more processors configured to: performing psychoacoustic audio decoding on the encoded foreground audio signal to obtain a foreground audio signal; determining a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding with respect to the encoded foreground audio signal; dequantizing the quantized spatial components to obtain scaled spatial components; descaling the scaled spatial component based on the bit allocation to the encoded foreground audio signal to obtain a spatial component; as well as The scene-based audio data is reconstructed based on the foreground audio signal and the spatial component.
12. The apparatus of claim 11, wherein: The one or more processors are configured to perform psychoacoustic audio decoding according to an AptX compression algorithm with respect to the encoded foreground audio signal to obtain the foreground audio signal.
13. The apparatus of claim 11, wherein: The one or more processors are configured to: obtaining, from the bitstream, a number of bits allocated to each of a coarse quantization gain and one or more fine quantization residuals, the coarse quantization gain and the one or more fine quantization residuals representing a gain of the foreground audio signal; as well as The scaled spatial component is descaled based on a number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain the spatial component.
14. The apparatus of claim 11, wherein: The scene-based audio data includes ambisonic reverberation coefficients corresponding to an order greater than one.
15. The apparatus of claim 11, wherein: The scene-based audio data includes ambisonic reverberation coefficients corresponding to orders greater than zero.
16. The apparatus of claim 11, wherein: The scene-based audio data includes audio data defined in a spherical harmonics domain.
17. The apparatus of claim 16, wherein: The encoded foreground audio signal comprises an encoded foreground audio signal defined in the spherical harmonics domain, and The scaled spatial components include scaled spatial components defined in the spherical harmonics domain.
18. The apparatus of claim 11, wherein: The one or more processors are further configured to: rendering the scene-based audio data to one or more speaker feeds; and A sound field represented by the scene-based audio data is reproduced based on the speaker feeds.
19. The apparatus of claim 11, wherein: The one or more processors are further configured to render the scene-based audio data to one or more speaker feeds, and Wherein the device comprises one or more loudspeakers configured to reproduce a sound field represented by the scene-based audio data based on the loudspeaker feeds.
20. The apparatus of claim 11, wherein: The scene-based audio data includes mixed-order ambisonic reverberation audio data.
21. A method of decoding a bitstream representing scene-based audio data, the method comprising: obtaining from the bitstream an encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of the encoded foreground audio signal; performing psychoacoustic audio decoding on the encoded foreground audio signal to obtain a foreground audio signal; determining a bit allocation for the encoded foreground audio signal when performing psychoacoustic audio decoding with respect to the encoded foreground audio signal; dequantizing the quantized spatial components to obtain scaled spatial components; descaling the scaled spatial component based on the bit allocation to the encoded foreground audio signal to obtain a spatial component; as well as The scene-based audio data is reconstructed based on the foreground audio signal and the spatial component.
22. The method of claim 21, wherein: Performing psychoacoustic audio decoding includes performing psychoacoustic audio decoding according to a compression algorithm with respect to the encoded foreground audio signal to obtain the foreground audio signal.
23. The method of claim 21, wherein: determining the bit allocation comprises obtaining from the bitstream a number of bits allocated to each of a coarse quantization gain and one or more fine quantization residuals, the coarse quantization gain and the one or more fine quantization residuals representing a gain of the foreground audio signal, and wherein descaling the scaled spatial component comprises descaling the scaled spatial component based on the number of bits allocated to each of the coarse quantization gain and the one or more fine quantization residuals to obtain the spatial component.
24. The method of claim 21, wherein: The scene-based audio data includes ambisonic reverberation coefficients corresponding to spherical basis functions having an order greater than zero.
25. The method of claim 21, wherein The scene-based audio data includes higher-order ambisonic reverberation coefficients corresponding to an order greater than one.
26. The method of claim 21, wherein: The scene-based audio data includes audio data defined in a spherical harmonics domain.
27. The method of claim 26, wherein: The encoded foreground audio signal comprises an encoded foreground audio signal defined in the spherical harmonics domain, and The scaled spatial components include scaled spatial components defined in the spherical harmonics domain.
28. The method of claim 21, further comprising: rendering the scene-based audio data to one or more speaker feeds; as well as A sound field represented by the scene-based audio data is reproduced based on the speaker feeds.
29. The method of claim 21, wherein: The scene-based audio data includes mixed-order ambisonic reverberation audio data.
30. A device configured to encode scene-based audio data, the device comprising: means for performing spatial audio coding with respect to the scene-based audio data to obtain a foreground audio signal and corresponding spatial components, the spatial components defining spatial characteristics of the foreground audio signal; means for performing psychoacoustic audio coding with respect to said foreground audio signal to obtain an encoded foreground audio signal; means for determining a bit allocation to the foreground audio signal when performing psychoacoustic audio coding with respect to the foreground audio signal; means for scaling the spatial component based on the bit allocation to the foreground audio signal to obtain a scaled spatial component; means for quantizing the scaled spatial components to obtain quantized spatial components; as well as means for specifying said encoded foreground audio signal and said quantized spatial components in a bitstream.
31. An apparatus configured to decode a bitstream representing encoded scene-based audio data, the apparatus comprising: means for obtaining from said bitstream an encoded foreground audio signal and corresponding quantized spatial components defining spatial characteristics of said encoded foreground audio signal; means for performing psychoacoustic audio decoding with respect to said encoded foreground audio signal to obtain a foreground audio signal; means for determining a bit allocation for said encoded foreground audio signal when performing psychoacoustic audio decoding with respect to said encoded foreground audio signal; means for dequantizing the quantized spatial components to obtain scaled spatial components; means for descaling the scaled spatial component based on the bit allocation to the encoded foreground audio signal to obtain a spatial component; as well as Means for reconstructing the scene-based audio data based on the foreground audio signal and the spatial component.
32. A computer-readable medium having one or more computer instructions recorded thereon, which, when executed by one or more processors of a device, cause the one or more processors to perform the method for encoding scene-based audio data according to claim 10.
33. A computer-readable medium having one or more computer instructions recorded thereon, which, when executed by one or more processors of a device, cause the one or more processors to perform the method for decoding a bitstream representing scene-based audio data according to any one of claims 21-29.
34. A computer program product comprising one or more computer instructions, which, when executed by one or more processors of a device, cause the one or more processors to perform the method for encoding scene-based audio data according to claim 10.
35. A computer program product comprising one or more computer instructions which, when executed by one or more processors of a device, cause the one or more processors to perform the method of decoding a bitstream representing scene-based audio data according to any one of claims 21-29.
Citation Information
Patent Citations
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US10405126B2
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US20190007781A1
Bitrate allocation for higher order ambisonic audio data
US10075802B1
Multi-stream audio coding
US20190103118A1