Correlating scene-based audio data for psychoacoustic audio coding
By using scene-based audio coding technology, which combines spatial audio coding and psychoacoustic coding, the sound field is separated into background and foreground signals, solving the problem of low encoding and decoding efficiency in existing technologies and achieving more efficient audio data transmission and storage.
Patent Information
- Application Number
- CN202080044737.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2020-06-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-06-23
AI Technical Summary
Existing technologies fail to effectively utilize the masking effect of the human auditory system when compressing audio data, resulting in low encoding and decoding efficiency.
Employing scene-based audio coding technology, this method combines spatial audio coding and psychoacoustic coding to separate the sound field into background and foreground audio signals. It then uses a psychoacoustic model for encoding and decoding to generate an efficient bitstream.
It improves the encoding and decoding efficiency of audio data, reduces data redundancy, and enhances the transmission and storage efficiency of audio data.
Smart Images

Figure CN114341976B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 16 / 908,032, filed June 22, 2020, entitled “CORRELATING SCENE-BASED AUDIODATA FOR PSYCHOACOUSTIC AUDIO CODING,” which claims priority to U.S. Provisional Application No. 62 / 865,865, filed June 24, 2019, entitled “CORRELATING SCENE-BASED AUDIODATA FOR PSYCHOACOUSTIC AUDIO CODING,” both of which are incorporated herein by reference in their entirety as set forth herein. Technical Field
[0002] This disclosure relates to audio data, and more specifically, to the encoding and decoding of audio data. Background Technology
[0003] Psychoacoustic audio encoding and decoding refers to the process of compressing audio data using psychoacoustic models. It leverages limitations of the human auditory system to compress audio data, taking into account constraints arising from spatial masking (e.g., two audio sources in the same location, where one source masks the other in terms of loudness), temporal masking (e.g., where one source masks the other in terms of loudness), and other factors. Psychoacoustic models attempt to model the human auditory system to identify redundancy, masking, or other portions of the sound field that are not perceptible to the human auditory system. Psychoacoustic audio encoding and decoding can also perform lossless compression by entropy encoding the audio data. Summary of the Invention
[0004] Typically, techniques for correlating scene-based audio data for use in psychoacoustic audio encoding and decoding are described.
[0005] In one example, aspects of the technology relate to a device configured to encode scene-based audio data, the device comprising: a memory configured to store the scene-based audio data; and one or more processors configured to: perform spatial audio coding on the scene-based audio data to obtain multiple background components, multiple foreground audio signals, and corresponding multiple spatial components of a sound field represented by the scene-based audio data, each of the multiple spatial components defining spatial characteristics of a corresponding foreground audio signal in the multiple foreground audio signals; perform correlation on two or more of the multiple background components and multiple foreground audio signals to obtain multiple correlated components; perform psychoacoustic audio coding on one or more of the multiple correlated components to obtain encoded components; and specify the encoded components in a bitstream.
[0006] In another example, aspects of the technology relate to a method for encoding scene-based audio data, the method comprising: performing spatial audio coding on the scene-based audio data to obtain a plurality of background components, a plurality of foreground audio signals, and a plurality of corresponding spatial components of a sound field represented by the scene-based audio data, each of the plurality of spatial components defining spatial characteristics of a corresponding foreground audio signal in the plurality of foreground audio signals; correlating the plurality of background components and one or more of the plurality of foreground audio signals to obtain a plurality of correlated components; performing psychoacoustic audio coding on one or more of the plurality of correlated components to obtain encoded components; and specifying the encoded components in a bitstream.
[0007] In another example, aspects of the technology relate to a device configured to encode scene-based audio data, the device comprising: means for performing spatial audio coding on the scene-based audio data to obtain a plurality of background components, a plurality of foreground audio signals, and a plurality of corresponding spatial components of a sound field represented by the scene-based audio data, each of the plurality of spatial components defining spatial characteristics of a corresponding foreground audio signal in the plurality of foreground audio signals; means for performing correlation on two or more of the plurality of background components and the plurality of foreground audio signals to obtain a plurality of correlated components; means for performing psychoacoustic audio coding on one or more of the plurality of correlated components to obtain encoded components; and means for specifying the encoded components in a bitstream.
[0008] In another example, aspects of the technology relate to a non-transitory computer-readable storage medium on which instructions are stored, which, when executed, cause one or more processors to: perform spatial audio coding on scene-based audio data to obtain multiple background components, multiple foreground audio signals, and corresponding multiple spatial components of a sound field represented by the scene-based audio data, each of the multiple spatial components defining spatial characteristics of a corresponding foreground audio signal among the multiple foreground audio signals; perform correlation on one or more of the multiple background components and multiple foreground audio signals to obtain multiple correlated components; perform psychoacoustic audio coding on one or more of the multiple correlated components to obtain encoded components; and specify the encoded components in a bitstream.
[0009] In another example, aspects of the technology relate to a device configured to decode a bitstream representing scene-based audio data, the device comprising: a memory configured to store the bitstream, the bitstream including multiple encoded correlated components of a sound field represented by the scene-based audio data; and one or more processors configured to: perform psychoacoustic audio decoding for one or more of the multiple encoded correlated components to obtain the multiple correlated components; obtain from the bitstream an indication of how one or more of the multiple correlated components are reordered in the bitstream; reorder the multiple correlated components based on the indication to obtain multiple reordered components; and reconstruct the scene-based audio data based on the multiple reordered components.
[0010] In another example, aspects of the technology relate to a method for decoding a bitstream representing scene-based audio data, the method including: obtaining multiple encoded related components from the bitstream; performing psychoacoustic audio decoding on one or more of the multiple encoded related components to obtain multiple related components; obtaining an indication from the bitstream indicating how one or more of the multiple related components are reordered in the bitstream; reordering the multiple related components based on the indication to obtain multiple reordered components; and reconstructing the scene-based audio data based on the multiple reordered components.
[0011] In another example, aspects of the technology relate to a device configured to decode a bitstream representing scene-based audio data, the device comprising: means for obtaining a plurality of encoded related components from the bitstream; means for performing psychoacoustic audio decoding on one or more of the plurality of encoded related components to obtain the plurality of related components; means for obtaining from the bitstream an indication of how one or more of the plurality of related components are reordered in the bitstream; means for reordering the plurality of related components based on the indication to obtain a plurality of reordered components; and means for reconstructing the scene-based audio data based on the plurality of reordered components.
[0012] In another example, aspects of the technology involve a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: obtain multiple encoded related components from a bitstream representing scene-based audio data; perform psychoacoustic audio decoding on one or more of the multiple encoded related components to obtain multiple related components; obtain from the bitstream an indication of how one or more of the multiple related components are reordered in the bitstream; reorder the multiple related components based on the indication to obtain multiple reordered components; and reconstruct the scene-based audio data based on the multiple reordered components.
[0013] Details of one or more aspects of the technology are set forth in the accompanying drawings and the following description. Other features, objects, and advantages of these technologies will be apparent from the description and drawings, as well as from the claims. Attached Figure Description
[0014] Figure 1 It is an illustration of a system that can perform various aspects of the techniques described in this disclosure.
[0015] Figure 2 This is an illustration of another example of a system that can perform various aspects of the techniques described in this disclosure.
[0016] Figures 3A-3C It's illustrated in more detail. Figure 1 and 2 The example shown is a block diagram of an example psychoacoustic audio encoding device.
[0017] Figure 4A and 4B It's illustrated in more detail. Figure 1 and 2 The example shown is a block diagram of an example psychoacoustic audio decoding device.
[0018] Figure 5 It's illustrated in more detail. Figures 3A-3C The example shown is a block diagram of an encoder.
[0019] Figure 6 It's illustrated in more detail. Figure 4A and Figure 4B A block diagram of an example decoder.
[0020] Figure 7 It's illustrated in more detail. Figures 3A-3C The example shown is a block diagram of an encoder.
[0021] Figure 8 It's illustrated in more detail. Figure 4A and Figure 4B The example shows a block diagram illustrating the implementation of the decoder.
[0022] Figure 9A and Figure 9B It's illustrated in more detail. Figures 3A-3C A block diagram of another example of the encoder shown in the example.
[0023] Figure 10A and Figure 10B It's illustrated in more detail. Figure 4A and Figure 4B A block diagram of another example of the decoder shown in the example.
[0024] Figure 11 This is an illustration of an example of top-down quantization.
[0025] Figure 12 This is an illustration of an example of bottom-up quantization.
[0026] Figure 13 It's a diagram. Figure 2 The example shown is a block diagram of an example component of the source device.
[0027] Figure 14 It's a diagram. Figure 2 A block diagram of exemplary components of a sink device is shown in the example.
[0028] Figure 15 It's a diagram. Figure 1 The example shown is a flowchart illustrating exemplary operation of an audio encoder when performing various aspects of the techniques described in this disclosure.
[0029] Figure 16 It's a diagram. Figure 1 The example shown is a flowchart illustrating exemplary operation of an audio decoder when performing various aspects of the techniques described in this disclosure. Detailed Implementation
[0030] There are different types of audio formats, including channel-based, object-based, and scene-based formats. Scene-based formats can use high-fidelity stereo technology. This high-fidelity stereo technology allows the sound field to be represented using a layered set of elements from speaker feeds that can be rendered to most speaker configurations.
[0031] An example of a hierarchical set of elements is the spherical harmonic coefficient (SHC) set. The following expression shows a description or representation of a sound field using SHCs:
[0032]
[0033] This expression shows the time t at any point in the sound field. Pressure p at the point i It can be made by SHC, Uniquely indicated. Here, c is the speed of sound (~343 m / s). It is the reference point (or observation point), j n (·) is an nth-order spherical Bessel function, and These are the nth and mth order spherical harmonic basis functions (also called spherical basis functions). It can be recognized that the terms within the square brackets are signals (i.e., The frequency domain representation of can be approximated by various time-frequency transforms, such as the Discrete Fourier Transform (DFT), Discrete Cosine Transform (DCT), or wavelet transform. Other examples of hierarchical sets include wavelet transform coefficient sets and other coefficient sets for multi-resolution basis functions.
[0034] It can be physically acquired (e.g., recorded) by various microphone array configurations, or alternatively, it can be derived from a channel-based or object-based description of the sound field (e.g., a pulse code modulation (PCM) audio object, which includes the audio object and metadata defining the location of the audio object within the sound field). SHC (which can also be called high-fidelity stereo sound coefficient) represents scene-based audio, where the SHC can be input to an audio encoder to obtain an encoded SHC that can facilitate more efficient transmission or storage. For example, it can use (1+4) 2 (25, therefore a fourth-order) coefficient fourth-order representation.
[0035] As mentioned above, SHC can be derived from microphone recordings using a microphone array. Various examples of how to derive SHC from a microphone array are described below: Poletti, M., “Three-Dimensional Surround Sound Systems Based on Spherical Harmonics”, J. Audio Eng. Soc., Vol. 53, No. 11, November 2005, pp. 1004-1025.
[0036] To illustrate how SHC is derived from an object-based description, consider the following equation: Coefficients corresponding to the sound field of a single audio object. It can be represented as:
[0037]
[0038] Where i is It is an n-order spherical Hankel function (of the second kind), and It refers to the location of the object. Knowing the object's source energy g(ω) as a function of frequency (e.g., using time-frequency analysis techniques, such as performing a Fast Fourier Transform on a PCM stream) allows for the conversion of each PCM object and its corresponding location into... Furthermore, it can be seen (since the above is a linear orthogonal decomposition) that for each object The coefficients are additive. In this way, multiple PCM objects (where a PCM object is an example of an audio object) can be generated by... The coefficients are represented (e.g., as the sum of coefficient vectors for each individual object). Essentially, these coefficients contain information about the sound field (pressure as a function of 3D coordinates), and are represented above at the viewpoint. The transformation from a single object to the representation of the entire sound field. The following figures are described in the context of SHC-based audio codecs.
[0039] Figure 1 This is a diagram illustrating a system 10 capable of performing various aspects of the techniques described in this disclosure. For example... Figure 1 As shown in the example, system 10 includes a content creator system 12 and a content consumer 14. Although described in the context of content creator system 12 and content consumer 14, the technique can be implemented in any context in which the SHC (which may also be referred to as high-fidelity stereo coefficients) or any other hierarchical representation of the sound field is encoded to form a bitstream representing audio data.
[0040] Furthermore, the content creator system 12 may represent a system comprising one or more of any form of computing device capable of implementing the techniques described in this disclosure, including any one or more of a handheld device (or cellular phone, including so-called "smartphones," or, in other words, a mobile phone or handheld device), a tablet computer, a laptop computer, a desktop computer, an extended reality (XR) device (which may refer to virtual reality—VR—devices, augmented reality—AR—devices, mixed reality—MR—devices, etc.), a gaming system, a CD player, a receiver (such as an audio / visual—A / V—receiver), or dedicated hardware to provide some examples.
[0041] Similarly, content consumer 14 may refer to any form of computing device capable of implementing the technology described in this disclosure, including handheld devices (or cellular phones, including so-called "smartphones," or in other words, mobile handheld devices or telephones), XR devices, tablet computers, televisions (including so-called "smart TVs"), set-top boxes, laptop computers, gaming systems or consoles, watches (including so-called smartwatches), wireless headsets (including so-called "smart headsets"), or desktop computers to provide several examples.
[0042] The content creator system 12 can represent any entity that can generate audio content and, possibly, video content for consumption by content consumers (such as content consumer 14). The content creator system 12 can capture live audio data from events such as sporting events, while also inserting various other types of additional audio data, such as commentary audio data, commercial audio data, introductory or exit audio data, into the live audio content.
[0043] Content consumer 14 refers to an individual who owns or has access to audio playback system 16, which can refer to any form of audio playback system capable of rendering high-order high-fidelity stereo audio data (including high-order audio coefficients, which may also be called spherical harmonic coefficients) to a speaker feed for playback as audio content. Figure 1 In the example, content consumer 14 includes audio playback system 16.
[0044] High-fidelity stereo audio data can be defined in the spherical harmonic function domain and rendered or otherwise transformed from the spherical harmonic function domain to the spatial domain, thereby producing audio content fed by one or more speakers. High-fidelity stereo audio data can represent an example of "scene-based audio data," which describes an audio scene using high-fidelity stereo coefficients. Scene-based audio data differs from object-based audio data in that, unlike the discrete objects (in the spatial domain) commonly found in object-based audio data, the entire scene is described (in the spherical harmonic function domain). Scene-based audio data also differs from channel-based audio data in that it resides in the spherical harmonic domain, which is the opposite of the spatial domain of channel-based audio data.
[0045] In any case, the content creator system 12 includes a microphone 18 that records or acquires live recordings in various formats, including directly as high-fidelity stereo sound coefficients and audio objects. When the microphone array 18 (which may also be referred to as "microphone 18") acquires live audio directly as high-fidelity stereo sound coefficients, the microphone 18 may include, for example... Figure 1 The example shows a high-fidelity stereo transcoder 20.
[0046] In other words, although shown as separate from microphone 5, separate instances of the high-fidelity stereo transcoder 20 can be included in each microphone 5 to transcode the acquired feed into high-fidelity stereo coefficients 21. However, when not included within microphone 18, the high-fidelity stereo transcoder 20 can transcode the live feed output from microphone 18 into high-fidelity stereo coefficients 21. In this respect, the high-fidelity stereo transcoder 20 can refer to a unit configured to transcode microphone feeds and / or audio objects into high-fidelity stereo coefficients 21. Therefore, the content creator system 12 includes a high-fidelity stereo transcoder 20 integrated with microphone 18, a high-fidelity stereo transcoder separate from microphone 18, or some combination thereof.
[0047] The content creator system 12 may also include an audio encoder 22 configured to compress high-fidelity stereo coefficients 21 to obtain a bitstream 31. The audio encoder 22 may include a spatial audio encoding device 24 and a psychoacoustic audio encoding device 26. The spatial audio encoding device 24 may represent a device capable of performing compression on the high-fidelity stereo coefficients 21 to obtain intermediate-formatted audio data 25 (which may also be referred to as “mezzanine-formatted audio data 25” when the content creator system 12 represents a broadcast network, as described in more detail below). The intermediate-formatted audio data 25 may represent audio data compressed using spatial audio compression but not yet subjected to psychoacoustic audio encoding (e.g., AptX or Advanced Audio Codec—AAC, or other similar types of psychoacoustic audio encoding, including various enhanced AAC—eAAC—such as High Efficiency AAC—HE-AAC—HE-AAC v2, also known as eAAC+, etc.).
[0048] Spatial audio coding device 24 can be configured to compress high-fidelity stereo coefficients 21. That is, spatial audio coding device 24 can use a decomposition involving the application of a linear invertible transform (LIT) to compress the high-fidelity stereo coefficients 21. An example of a linear invertible transform is referred to as “singular value decomposition” (“SVD”), principal component analysis (“PCA”), or eigenvalue decomposition, which can represent different examples of linear invertible decomposition.
[0049] In this example, the spatial audio coding device 24 can apply SVD to the high-fidelity stereo coefficients 21 to determine a decomposed version of the high-fidelity stereo coefficients 21. The decomposed version of the high-fidelity stereo coefficients 21 may include one or more primary audio signals and one or more corresponding spatial components that describe the spatial characteristics of the associated primary audio signals, such as orientation, shape, and width. Therefore, the spatial audio coding device 24 can apply decomposition to the high-fidelity stereo coefficients 21 to decouple energy (as represented by the primary audio signals) from spatial characteristics (as represented by the spatial components).
[0050] Spatial audio coding device 24 can analyze the decomposed versions of high-fidelity stereo coefficients 21 to identify various parameters, which can facilitate the reordering of the decomposed versions of high-fidelity stereo coefficients 21. Spatial audio coding device 24 can reorder the decomposed versions of high-fidelity stereo coefficients 21 based on the identified parameters, where this reordering can improve encoding / decoding efficiency, provided that the reordering of high-fidelity stereo coefficients is performed across frames that span the high-fidelity stereo coefficients (where a frame typically includes M samples of the decomposed versions of high-fidelity stereo coefficients 21, and in some examples, M is set to 1024).
[0051] After reordering the decomposed versions of the high-fidelity stereo coefficients 21, the spatial audio coding device 24 can select one or more of the decomposed versions of the high-fidelity stereo coefficients 21 as a representation of the foreground (or in other words, the prominent, dominant, or salient) component of the sound field. The spatial audio coding device 24 can specify the decomposed version of the high-fidelity stereo coefficients 21 representing the foreground component (which may also be called the "primary sound signal," "primary audio signal," or "primary sound component") and the associated directional information (which may also be called the "spatial component," or in some cases, the so-called "V vector" that identifies the spatial characteristics of the corresponding audio object). The spatial component can represent a vector with multiple distinct elements (which, in the context of a vector, may be called a "coefficient"), and therefore can be called a "multidimensional vector."
[0052] Spatial audio encoding device 24 can then perform sound field analysis on the high-fidelity stereo coefficients 21 to at least partially identify the high-fidelity stereo coefficients 21 representing one or more background (or in other words, ambient) components of the sound field. The background components may also be referred to as “background audio signals” or “ambient audio signals.” In some examples, assuming the background audio signal may consist only of a subset of any given samples of the high-fidelity stereo coefficients 21 (e.g., those corresponding to zero-order and first-order spherical basis functions, but not those corresponding to second-order or higher-order spherical basis functions), then spatial audio encoding device 24 can perform energy compensation on the background audio signal. In other words, when performing order reduction, spatial audio encoding device 24 can increase the residual background high-fidelity stereo coefficients 21 (e.g., add energy to them / subtract energy from them) to compensate for the change in total energy caused by performing order reduction.
[0053] Spatial audio encoding device 24 may then perform a form of interpolation (another way of referring to spatial components) on the foreground direction information, and then perform order reduction on the interpolated foreground direction information to produce reduced foreground direction information. In some examples, spatial audio encoding device 24 may also perform quantization on the reduced foreground direction information to output encoded foreground direction information. In some cases, this quantization may include scalar / entropy quantization, which may be in the form of vector quantization. Spatial audio encoding device 24 may then output the intermediately formatted audio data 25 as a background audio signal, a foreground audio signal, and quantized foreground direction information.
[0054] In any case, in some examples, the background audio signal and the foreground audio signal may include a transport channel. That is, the spatial audio coding device 24 may output a transport channel for each frame of the high-fidelity stereo coefficients 21 of the corresponding one in the background audio signal (e.g., M samples of one of the high-fidelity stereo coefficients 21 corresponding to a zeroth or first-order spherical basis function) and for each frame of the foreground audio signal (e.g., M samples of the audio object decomposed from the high-fidelity stereo coefficients 21). The spatial audio coding device 24 may also output side information (which may also be referred to as “sideband information”), which includes the quantized spatial components corresponding to each of the foreground audio signals.
[0055] In general, Figure 1 In the example, the transmission channel and auxiliary information can be represented as High Fidelity Stereo Transmission Format (ATF) audio data 25 (this refers to another way of formatting audio data). In other words, AFT audio data 25 can include transmission channel and auxiliary information (also known as "metadata"). As an example, ATF audio data 25 can conform to HOA (High-Order High Fidelity Stereo) transmission format (HTF). More information on HTF can be found in the European Telecommunications Standards Institute (ETSI) technical specification (TS) entitled "High-Order High Fidelity Stereo (HOA) Transmission Format" ETSI TS 103 589 V1.1.1, dated June 2018 (2018-06). Thus, ATF audio data 25 can be referred to as HTF audio data 25.
[0056] Then, spatial audio coding device 24 can transmit or output ATF audio data 25 to psychoacoustic audio coding device 26. Psychoacoustic audio coding device 26 can perform psychoacoustic audio coding on the ATF audio data 25 to generate a bitstream 31. Psychoacoustic audio coding device 26 can operate according to standardized, open-source, or proprietary audio encoding / decoding processes. For example, psychoacoustic audio coding device 26 can perform psychoacoustic audio coding according to any type of compression algorithm, such as those provided by the Moving Picture Experts Group (MPEG), the MPEG-H 3D audio coding standard, the MPEG-I immersive audio standard, or proprietary standards (such as AptX). TM(Including various versions of AptX, such as Enhanced AptX—E-AptX, AptX Live, AptX Stereo, and AptX High Definition—AptX-HD), Advanced Audio Coding (AAC), Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey's Audios, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus, and Media Audio (WMA)) The unified speech and audio codecs, denoted as "USAC," are described. The content creator system 12 can then send the bitstream 31 to the content consumer 14 via the transport channel.
[0057] In some examples, the psychoacoustic audio encoding device 26 may represent one or more examples of psychoacoustic audio codecs, each of which is used for encoding the transmission channel of the ATF audio data 25. In some cases, the psychoacoustic audio encoding device 26 may represent one or more examples of AptX encoding units (as described above). In some cases, the psychoacoustic audio codec unit 26 may invoke an example of a stereo encoding unit for each transmission channel of the ATF audio data 25.
[0058] In some examples, in order to generate different representations of the sound field using high-fidelity stereo coefficients (which is again an example of audio data 21), audio encoder 22 may use an encoding and decoding scheme for the high-fidelity stereo representation of the sound field, referred to as Mixed-Order Ambison (MOA), as described in detail below: U.S. Application Serial No. 15 / 672,058, filed August 8, 2017, entitled “MIXED-ORDER ambisonics (MOA) AUDIO DATA FO COMPUTER-MEDIATED REALITY SYSTEMS”, and published January 3, 2019, as U.S. Patent Publication No. 2019 / 0007781.
[0059] To generate a specific MOA representation of a sound field, the audio encoder 22 can generate a subset of the entire set of high-fidelity stereo coefficients. For example, each MOA representation generated by the audio encoder 22 may provide accuracy for some regions of the sound field, but lower accuracy for others. In one example, the MOA representation of the sound field may include eight (8) uncompressed high-fidelity stereo coefficients, while a third-order high-fidelity stereo representation of the same sound field may include sixteen (16) uncompressed high-fidelity stereo coefficients. Thus, each MOA representation of the sound field generated as a subset of the high-fidelity stereo coefficients may have lower storage density and lower bandwidth density than a corresponding third-order high-fidelity stereo representation of the same sound field generated from the high-fidelity stereo coefficients (if and when transmitted as part of a bitstream 31 on the illustrated transport channel).
[0060] Although described in relation to MOA representation, the techniques of this disclosure can also be performed for a first-order high-fidelity stereo (FOA) representation, wherein all high-fidelity stereo coefficients corresponding to spherical basis functions of order up to one are used to represent the sound field. In other words, the sound field representation generator 302 can represent the sound field using all high-fidelity stereo coefficients of order one, rather than using a subset of non-zero high-fidelity stereo coefficients.
[0061] In this regard, high-order high-fidelity stereo audio data may include high-order high-fidelity stereo coefficients associated with spherical basis functions of one or less order (which may be referred to as "first-order high-fidelity stereo audio data"), high-order high-fidelity stereo coefficients associated with spherical basis functions of mixed order and sub-order (which may be referred to as "MOA representation" discussed above), or high-order high-fidelity stereo coefficients associated with spherical basis functions of more than one order.
[0062] In addition, although Figure 1The bitstream 31 is shown as being sent directly to content consumer 14, but content creator system 12 can output the bitstream 31 to an intermediate device located between content creator system 12 and content consumer 14. The intermediate device can store the bitstream 31 for later delivery to content consumer 14, which can request the bitstream. The intermediate device can include a file server, web server, desktop computer, laptop computer, tablet computer, mobile phone, smartphone, or any other device capable of storing the bitstream 31 for later retrieval by an audio decoder. The intermediate device can reside in a content delivery network capable of streaming the bitstream 31 (and possibly in conjunction with sending corresponding video data bitstreams) to subscribers requesting the bitstream 31, such as content consumer 14.
[0063] Alternatively, the content creator system 12 may store the bitstream 31 to a storage medium, such as an optical disc, digital video disc, high-definition video disc, or other storage medium, most of which is computer-readable and therefore may be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this context, a transmission channel may refer to those channels (and may include retail stores and other store-based delivery mechanisms) used to transmit content stored on these media. In any case, the technology disclosed herein should therefore not be limited in this respect. Figure 1 Examples.
[0064] For example Figure 1 As shown in the example, content consumer 14 includes audio playback system 16. Audio playback system 16 can represent any audio playback system capable of playing back multichannel audio data. Audio playback system 16 may also include audio decoding device 32. Audio decoding device 32 can represent a device configured to decode high-fidelity stereo coefficients 11' from bitstream 31, wherein high-fidelity stereo coefficients 11' may be similar to high-fidelity stereo coefficients 11 but differ due to lossy operation (e.g., quantization) and / or transmission via a transmission channel.
[0065] Audio decoding device 32 may include psychoacoustic audio decoding device 34 and spatial audio decoding device 36. Psychoacoustic audio decoding device 34 may represent a unit configured to operate inversely to psychoacoustic audio encoding device 26 to reconstruct ATF audio data 25' from bitstream 31. Furthermore, the primary notation for ATF audio data 25' output from psychoacoustic audio decoding device 34 may differ slightly from ATF audio data 25 due to lossy or other operations performed during compression of ATF audio data 25. Psychoacoustic audio decoding device 34 may be configured to perform decompression according to standardized, open-source, or proprietary audio codec processes (such as AptX, variants of AptX, AAC, variants of AAC, etc., as described above).
[0066] Although the following description focuses primarily on AptX, this technique can be applied to other psychoacoustic audio codecs. Examples of other psychoacoustic audio codecs include Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), and MPEG-4 Lossless Audio Streaming (ALS). Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey's Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus, and Windows Media Audio (WMA).
[0067] In any case, the psychoacoustic audio decoding device 34 can perform psychoacoustic decoding on the foreground audio object specified in the bitstream 31 and the encoded high-fidelity stereo sound coefficients representing the background audio signal specified in the bitstream 31. In this way, the psychoacoustic audio decoding device 34 can obtain ATF audio data 25' and output the ATF audio data 25' to the spatial audio decoding device 36.
[0068] Spatial audio decoding device 36 can represent a unit configured to operate inversely to spatial audio encoding device 24. That is, spatial audio decoding device 36 can dequantize specified foreground direction information in bitstream 31. Spatial audio decoding device 36 can also perform dequantization on the quantized foreground direction information to obtain decoded foreground direction information. Spatial audio decoding device 36 can then perform interpolation on the decoded foreground direction information, and then determine high-fidelity stereo coefficients representing the foreground component based on the decoded foreground audio signal and the interpolated foreground direction information. Spatial audio decoding device 36 can then determine high-fidelity stereo coefficients 11' based on the determined high-fidelity stereo coefficients representing the foreground audio signal and the decoded high-fidelity stereo coefficients representing the background audio signal.
[0069] After decoding the bitstream 31 to obtain the high-fidelity stereo response coefficient 11', the audio playback system 16 can render the high-fidelity stereo response coefficient 11' to the output speaker feed 39. The audio playback system 16 may include multiple different audio renderers 38. Each audio renderer 38 may provide different forms of rendering, which may include one or more of various methods of performing vector basis amplitude shift (VBAP), one or more of various methods of performing binaural rendering (e.g., head-related transfer function—HRTF, binaural room impulse response—BRIR, etc.), and / or one or more of various methods of performing sound field synthesis.
[0070] The audio playback system 16 can output a speaker feed 39 to one or more speakers 40. The speaker feed 39 can drive the speakers 40. The speakers 40 can represent loudspeakers (e.g., transducers placed in a cabinet or other enclosure), headphone speakers, or any other type of transducer capable of emitting sound based on electrical signals.
[0071] To select an appropriate renderer, or in some cases generate an appropriate renderer, the audio playback system 16 may obtain speaker information 41 indicating the number of speakers 40 and / or the spatial geometry of the speakers 40. In some cases, the audio playback system 16 may use a reference microphone to obtain the speaker information 41 and drive the speakers 40 in a manner that dynamically determines the speaker information 41. In other cases, or in conjunction with the dynamic determination of the speaker information 41, the audio playback system 16 may prompt the user to interface with the audio playback system 16 and input the speaker information 41.
[0072] The audio playback system 16 can select one of the audio renderers 38 based on the speaker information 41. In some cases, the audio playback system 16 can generate one of the audio renderers 38 based on the speaker information 41 when no audio renderer 38 is within the threshold similarity metric (according to speaker geometry) specified in the speaker information 41. In some cases, the audio playback system 16 can generate one of the audio renderers 38 based on the speaker information 41 without first attempting to select one of the existing audio renderers 38.
[0073] Although speaker feed 39 has been described, audio playback system 16 can render headphone feeds from speaker feed 39 or directly from high-fidelity stereo sound coefficient 11', thereby outputting the headphone feeds to headphone speakers. The headphone feeds can represent binaural audio speaker feeds, which audio playback system 16 renders using a binaural audio renderer. As described above, audio encoder 22 can invoke spatial audio encoding device 24 to perform spatial audio encoding (or otherwise compress) on high-fidelity stereo audio data 21, thereby obtaining ATF audio data 25. During the application of spatial audio encoding to high-fidelity stereo audio data 21, spatial audio encoding device 24 can obtain the foreground audio signal and the corresponding spatial components, which are respectively designated as transport channels and accompanying metadata (or sideband information) in decoded form.
[0074] As described above, the spatial audio coding device 24 can encode the high-fidelity stereo audio data 21 to obtain ATF audio data 25, which may include multiple transmission channels specifying multiple background components, multiple foreground audio signals, and corresponding multiple spatial components. In some examples, when conforming to HTF, the ATF audio data 25 may include four foreground audio signals and a first-order high-fidelity stereo audio signal, whose coefficients corresponding to both the zeroth-order spherical basis function and three first-order spherical basis functions serve as background components, for a total of four background components. The spatial audio coding device 24 can output the ATF audio data 25 to the psychoacoustic audio coding device 26, which can perform some form of psychoacoustic audio coding.
[0075] In some examples, the psychoacoustic audio coding device 26 may perform a form of stereo psychoacoustic audio coding and decoding, wherein prediction is performed between at least two transmission channels of the ATF audio data 25 to determine differences, thereby potentially reducing the dynamic range of the transmission channels. Assuming the stereo audio data includes only two channels that are relatively correlated in height and position (relatively, although their phases may differ), the stereo psychoacoustic audio coding algorithm may not perform any correlation for the stereo audio data. Thus, applying stereo psychoacoustic audio coding to any given pair of transmission channels of the ATF audio data 25 may result in lower compression efficiency, because any given pair of transmission channels may or may not have sufficient correlation to achieve adequate dynamic gain reduction.
[0076] According to various aspects of the technology described in this disclosure, the psychoacoustic audio coding apparatus 26 may perform correlation for two or more transmission channels of the ATF audio data 25 to obtain correlation values before performing stereo or multichannel psychoacoustic audio coding for the transmission channels. In some examples, the psychoacoustic audio coding apparatus 26 may perform correlation for each unique transmission channel pair specifying a background component and for each unique transmission channel pair specifying a foreground audio signal. In some examples, the psychoacoustic audio coding apparatus 26 may perform correlation for each transmission channel unit pair (as an example, comparing at least one background component with at least one foreground audio signal).
[0077] In any case, the psychoacoustic audio coding device 26 can then reorder the transport channels based on correlation values (pairing transport channels according to the highest correlation value). By potentially improving the correlation before stereo psychoacoustic audio coding, this technique can potentially improve coding and decoding efficiency, thereby improving the operation of the psychoacoustic audio coding device 26 itself.
[0078] In operation, the spatial audio encoding device 24 can perform spatial audio encoding on the scene-based audio data 21 to obtain multiple background components, multiple foreground audio signals, and corresponding multiple spatial components as ATF audio data 25. The spatial audio encoding device 24 can output the ATF audio data 25 to the psychoacoustic audio encoding device 26.
[0079] The psychoacoustic audio encoding device 26 can receive background and foreground audio signals. For example... Figure 1 As illustrated in the example, the psychoacoustic audio coding device 26 may include a correlation unit (CU) 46, which can perform the aforementioned correlation with two or more of a plurality of background components and a plurality of foreground audio signals to obtain a plurality of correlated components. As described above, the CU 46 may perform correlation with the background components and the foreground audio signals separately. In other examples, as described above, the CU 46 may perform correlation with both the background components and the foreground audio signals, wherein at least one background component and at least one foreground audio signal undergo correlation.
[0080] CU 46 can obtain correlation values as a result of performing correlation on two or more background components and foreground audio signals. CU 46 can reorder the background components and foreground audio signals based on the correlation values. CU 46 can output reordering metadata indicating how the transport channel is reordered to spatial audio coding device 24, which can specify the reordering metadata in metadata including spatial components. Although described as specifying reordering metadata in metadata including spatial components, psychoacoustic audio coding device 26 can specify the reordering metadata in bitstream 31.
[0081] In any case, CU 46 can output a reordered transmission channel (which can specify multiple correlated components), so that psychoacoustic audio coding device 26 can perform psychoacoustic audio coding on the multiple correlated components to obtain encoded components. As described in more detail below, psychoacoustic audio coding device 26 can perform psychoacoustic audio coding on at least one pair of the multiple correlated components according to the AptX compression algorithm to obtain multiple encoded components. Psychoacoustic audio coding device 26 can specify multiple encoded components in bitstream 31.
[0082] As described above, the audio decoder 32 is interoperable with the audio encoder 22. Thus, the audio decoder 32 can obtain the bitstream 31 and invoke the psychoacoustic audio decoding device 34. As described above, the psychoacoustic audio decoding device 34 can perform psychoacoustic audio decoding according to the AptX decompression algorithm. More information regarding the AptX decompression algorithm is also available in the references below. Figures 5-1 Example description of 0.
[0083] In any case, the psychoacoustic audio decoding device 34 can obtain reordering metadata from the bitstream 31, which can indicate how one or more of a plurality of related components are reordered in the bitstream 31. Figure 1 As illustrated in the example, the psychoacoustic audio decoding device 34 may include a reordering unit (RU) 54, which represents a unit configured to reorder multiple related components based on reordering metadata to obtain multiple reordered components. The psychoacoustic audio decoding device 34 may reconstruct ATF audio data 25' based on the multiple reordered components. The spatial audio decoding device 36 may then reconstruct scene-based audio data 21' based on the ATF audio data 25'.
[0084] Figure 2 This is a diagram illustrating another example of a system that can perform various aspects of the techniques described in this disclosure. Figure 2 System 110 can represent Figure 1 An example of system 10 is shown in the example. Figure 2 As shown in the example, system 110 includes source device 112 and sink device 114, where source device 112 may represent an example of content creator system 12, and sink device 114 may represent an example of content consumer 14 and / or audio playback system 16.
[0085] Although source device 112 and sink device 114 have been described, in some cases source device 112 may operate as sink device, while in these and other cases sink device 114 may operate as source device. Therefore, Figure 2The example of system 110 shown is merely one example illustrating various aspects of the technology described in this disclosure.
[0086] In any case, as mentioned above, source device 112 may represent any form of computing device capable of implementing the technology described in this disclosure, including handheld devices (or cellular phones, including so-called "smartphones"), tablet computers, so-called smartphones, remote-controlled aircraft (such as so-called "drones"), robots, desktop computers, receivers (such as audio / visual AV receivers), set-top boxes, televisions (including so-called "smart TVs"), media players (such as digital video disc players, streaming media players, Blu-ray disc players, etc.), or any other device capable of wirelessly transmitting audio data to the junction device via a personal area network (PAN). For illustrative purposes, source device 112 is assumed to represent a smartphone.
[0087] The Meeting Point device 114 can represent any form of computing device capable of implementing the technology described in this disclosure, including handheld devices (or, in other words, cellular phones, mobile handheld devices, etc.), tablet computers, smartphones, desktop computers, wireless headsets (which may include wireless headsets with or without microphones, and so-called smart wireless headsets that include additional functions such as fitness monitoring, onboard music storage and / or playback, dedicated cellular capabilities, etc.), wireless speakers (including so-called "smart speakers"), watches (including so-called "smartwatches"), or any other device capable of reproducing a sound field based on audio data wirelessly transmitted via a PAN. Furthermore, for illustrative purposes, it is assumed that the Meeting Point device 114 represents a wireless headset.
[0088] like Figure 2 As shown in the example, source device 112 includes one or more applications (“apps”) 118A-118N (“app118”), mixing unit 120, audio encoder 122 (which includes a spatial audio encoding device—SAED—124 and a psychoacoustic audio encoding device—PAED—126), and wireless connectivity manager 128. Although in Figure 2 The example is not shown, but source device 112 may include a number of other elements that support the operation of application 118, including an operating system, various hardware and / or software interfaces (such as user interfaces, including graphical user interfaces), one or more processors, memory, storage devices, etc.
[0089] Each application 118 represents software (such as an instruction set stored on a non-transitory computer-readable medium) that configures system 110 to provide some functionality when executed by one or more processors of source device 112. To list a few examples, application 118 may provide messaging functionality (such as access to email, text messaging, and / or video messaging), voice calling functionality, video conferencing functionality, calendar functionality, audio streaming functionality, orientation functionality, mapping functionality, and gaming functionality. Application 118 may be a first-party application designed and developed by the same company that designs and sells the operating system executed by source device 112 (and is typically pre-installed on source device 112), or a third-party application accessible via a so-called "app store" or possibly pre-installed on source device 112. When executed, each application 118 may output audio data 119A-119N ("Audio Data 119") respectively.
[0090] In some examples, audio data 119 may be obtained from a microphone (not depicted, but similar) connected to the source device 112. Figure 1 The example shown uses microphone 5). Audio data 119 can include data similar to that described above for... Figure 1 The example discusses the high-fidelity stereo audio data 21, which has high-fidelity stereo coefficients, and this high-fidelity stereo audio data can be referred to as "scene-based audio data". Thus, audio data 119 can also be referred to as "scene-based audio data 119" or "high-fidelity stereo audio data 119".
[0091] Although described in relation to high-fidelity stereo audio data, the technique can be performed on high-fidelity stereo audio data that does not necessarily include coefficients corresponding to so-called “higher-order” spherical basis functions (e.g., spherical basis functions with an order greater than one). Therefore, the technique can be performed on high-fidelity stereo audio data that includes coefficients corresponding only to zero-order spherical basis functions or only to zero-order and first-order spherical basis functions.
[0092] Mixing unit 120 refers to a unit configured to mix one or more of the audio data 119 output by application 118 (as well as other audio data output by the operating system—such as alarms or other tones, including keypad press tones, ringtones, etc.) to generate mixed audio data 121. Audio mixing can refer to the process of combining multiple sounds (as described in audio data 119) into one or more channels. During mixing, mixing unit 120 may also manipulate and / or enhance the volume level (also referred to as “gain level”), frequency content, and / or panoramic position of the high-fidelity stereo audio data 119. In the context of streaming high-fidelity stereo audio data 119 over a wireless PAN session, mixing unit 120 may output mixed audio data 121 to audio encoder 122.
[0093] The audio encoder 122 may be similar to (if not substantially similar to) the above. Figure 1 The example describes an audio encoder 22. That is, an audio encoder 122 can represent a unit configured to encode mixed audio data 121 and thereby obtain encoded audio data in the form of a bitstream 131. In some examples, the audio encoder 122 can encode individual audio data within the audio data 119.
[0094] For illustrative purposes, refer to an example of the PAN protocol. It offers many different types of audio codecs (it is a word derived from combining the words "encode" and "decode"), and is scalable to include vendor-specific audio codecs. The Advanced Audio Distribution Profile (A2DP) indicates that support for A2DP requires support for the sub-band codecs specified in A2DP. A2DP also supports codecs proposed in MPEG-1 Part 3 (MP2), MPEG-2 Part 3 (MP3), MPEG-2 Part 7 (Advanced Audio Coding – AAC), MPEG-4 Part 3 (High Efficiency AAC – HE-AAC), and Adaptive Transform Acoustic Codec (ATRAC). Furthermore, as mentioned above, A2DP supports vendor-specific codecs, such as aptX. TM And various other versions of aptX (e.g., enhanced aptX – E-aptX, aptXlive, and aptX High Definition – aptX-HD).
[0095] Audio encoder 122 can operate uniformly with any of the audio codecs listed above, as well as one or more audio codecs not listed above, but operates to encode mixed audio data 121 to obtain encoded audio data 131 (this refers to another way of referring to bitstream 131). Audio encoder 122 may first invoke SAED 124, which can be used with... Figure 1 The example shown is similar to (if not substantially similar to) SAED 24. SAED 124 can perform the above spatial audio compression on the mixed audio data to obtain ATF audio data 125 (if not similar to...). Figure 1 The example shown is essentially similar to ATF audio data 25, so it can be similar. SAED 124 can output ATF audio data 25 to PAED 126.
[0096] PAED 126 can be used with Figure 1 The example shows a PAED 26 similar to (if not substantially similar to) that. PAED 126 can perform psychoacoustic audio encoding according to any of the aforementioned codecs (including AptX and its variants) to obtain bitstream 131. Audio encoder 122 can output the encoded audio data 131 to one of the wireless communication units 130 (e.g., wireless communication unit 130A) managed by wireless connection manager 128.
[0097] The wireless connection manager 128 can represent units configured to allocate bandwidth within certain frequencies of the available spectrum to different wireless communication units 130. For example, The communication protocol operates within a 2.5 GHz spectrum range, which overlaps with the spectrum range used by various WLAN communication protocols. The wireless connection manager 128 can allocate a portion of the bandwidth to different devices during a given period. The protocol allocates different portions of bandwidth to overlapping WLAN protocols at different times. Bandwidth and other allocations are defined by scheme 129. Wireless connection manager 128 can expose various application programmer interfaces (APIs) through which the allocation of bandwidth and other aspects of the communication protocol can be adjusted to achieve a specified Quality of Service (QoS). In other words, wireless connection manager 128 can provide APIs to adjust scheme 129, through which scheme 129 controls the operation of wireless communication unit 130 to achieve the specified QoS.
[0098] In other words, the wireless connection manager 128 can manage the coexistence of multiple wireless communication units 130 operating within the same spectrum, such as certain WLAN communication protocols and some PAN protocols discussed above. The wireless connection manager 128 may include a coexistence scheme 129 (in... Figure 2As shown in the diagram as "Scheme 129", it indicates when (e.g., interval) each wireless communication unit 130 can send how many packets, the size of the packets sent, etc.
[0099] The wireless communication unit 130 may each represent a wireless communication unit 130 that operates according to one or more communication protocols to transmit bit stream 131 to the sink device 114 via a transmission channel. Figure 2 In the example, for illustrative purposes, it is assumed that the wireless communication unit 130A is based on The kit's communication protocol operation. It is also assumed that the wireless communication unit 130A operates according to A2DP to establish a PAN link (on the transport channel) to allow the bit stream 131 to be transferred from the source device 112 to the sink device 114. Although described in reference to a PAN link, this can be applied to cellular connections (such as so-called 3G, 4G, and / or 5G cellular data services), WiFi, etc. TM Any type of wired or wireless connection can be used to implement various aspects of this technology.
[0100] More information regarding the communication protocol suite can be found in the document titled "BluetoothCore Specification v 5.0," published on December 6, 2016, and is also available in the following literature: www.bluetooth.org / en-us / specification / adopted-specifications More information about A2DP can be found in the document titled "Advanced Audio Distribution Profile Specification" version 1.3.1, published on July 14, 2015.
[0101] The wireless communication unit 130A can output bit stream 131 to the sink device 114 via a transmission channel, which is assumed to be a wireless channel in the Bluetooth example. Although in Figure 2 The data is shown as being sent directly to sink device 114, but source device 112 can output bitstream 131 to an intermediate device located between source device 112 and sink device 114. The intermediate device can store bitstream 131 for later transmission to sink device 114, which can request bitstream 131. The intermediate device can include a file server, web server, desktop computer, laptop computer, tablet computer, mobile phone, smartphone, or any other device capable of storing bitstream 131 for later retrieval by an audio decoder. The intermediate device can reside in a content delivery network capable of streaming bitstream 131 (and possibly in conjunction with sending corresponding video data bitstreams) to subscribers requesting bitstream 131, such as sink device 114.
[0102] Alternatively, source device 112 may store bitstream 131 to a storage medium, such as an optical disc, digital video disc, high-definition video disc, or other storage medium, most of which is computer-readable and therefore may be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this context, a transmission channel may refer to those channels (and may include retail stores and other store-based delivery mechanisms) used to transmit content stored on these media. In any case, the technology disclosed herein should not therefore be limited in this respect. Figure 2 Examples.
[0103] like Figure 2 The example also shows that the junction device 114 includes: a wireless connection manager 150, which manages one or more of the wireless communication units 152A-152N (“Wireless Communication Units 152”) according to scheme 151; an audio decoder 132 (including a psychoacoustic audio decoding device—PADD—134 and a spatial audio decoding device—SADD—136); and one or more speakers 140A-140N (“Speaker 140”, which may be similar to Figure 1 (Speaker 40 shown in the example). The wireless connection manager 150 can operate in a manner similar to that described above for the wireless connection manager 128, exposing APIs to adjust scheme 151, through which the operation of the wireless communication unit 152 achieves the specified QoS.
[0104] Wireless communication unit 152 may be operationally similar to wireless communication unit 130, except that wireless communication unit 152 interoperates with wireless communication unit 130 to receive bit stream 131 via a transmission channel. It is assumed that one of the wireless communication units 152 (e.g., wireless communication unit 152A) according to... The kit operates on a communication protocol and is interchangeable with wireless communication protocols. The wireless communication unit 152A can output bitstream 131 to audio decoder 132.
[0105] Audio decoder 132 can operate in a manner reciprocal to audio encoder 122. Audio decoder 132 can operate consistently with any or more of the audio codecs listed above and those not listed above, but its operation is to decode the encoded audio data 131 to obtain mixed audio data 121'. Similarly, the primary designation for "mixed audio data 121" indicates that some loss may exist due to quantization or other lossy operations that occur during encoding by audio encoder 122.
[0106] Audio decoder 132 can call PADD 134 to perform psychoacoustic audio decoding on bitstream 131 to obtain ATF audio data 125', which PADD 134 can output to SADD 136. SADD 136 can perform spatial audio decoding to obtain mixed audio data 121'. Although for ease of illustration, Figure 2 The renderer is not shown in the example (similar to...) Figure 1 The renderer 38), but the audio decoder 132 can render the mixed audio data 121' to the speaker feed (using any renderer, such as the one described above for...). Figure 1 The example discussed is renderer 38) and the speaker feed output is sent to one or more of the speakers 140.
[0107] Each loudspeaker 140 represents a transducer configured to reproduce the sound field fed from the loudspeakers. The transducers can be as follows: Figure 2 The example shown is integrated within the sink device 114, or can be communicatively coupled to the sink device 114 (in a wired or wireless manner). Speaker 140 can represent any form of speaker, such as a loudspeaker, headphone speaker, or speaker in an earbud. Furthermore, although described in relation to a transducer, speaker 140 can represent other forms of speakers, such as the “loudspeaker” used in bone conduction headphones that transmit vibrations to the maxilla, which induce sound in the human auditory system.
[0108] As described above, PAED 126 can perform various aspects of the quantization technique described above for PAED 26 to quantize the spatial components based on the foreground audio signal correlated bit allocation. PADD 134 can also perform various aspects of the quantization technique described above for PADD 34 to dequantize the quantized spatial components based on the foreground audio signal correlated bit allocation. Figures 3A-3C The examples provide more information for PAED 126, while for PAED 126... Figure 4A and Figure 4B The example provides more information for PADD 134.
[0109] Figures 3A-3C It's illustrated in more detail. Figure 1 and Figure 2 The example shown is a block diagram of an example psychoacoustic audio encoding device. First refer to... Figure 3AFor example, the psychoacoustic audio encoder 226A can represent an example of PAED 26 and / or PAED 126. PAED 226A can receive transport channels 225A-225N from ATF encoder 224 (where the ATF encoder can represent another way of referring to spatial audio coding device 24). ATF encoder 224 can perform spatial audio coding for high-fidelity stereo coefficient 221 (which can represent an example of high-fidelity stereo coefficient 21), as described above for spatial audio coding device 24.
[0110] PAED 226A's CU 46A (which can represent) Figure 1 As shown in the example of CU 46, transmission channels 225 can be obtained, where transmission channels 225A-225D can specify the background component (and thus can be called background transmission channels 225), and transmission channels 225E-225H can specify the foreground audio signal (and thus can be called foreground transmission channels 225). Figure 3A As shown in the example, CU 46A may include a background (BG) related unit 228A, a foreground (FG) related unit 228B, a BG reordering unit 230A, and an FG reordering unit 230B.
[0111] BG related unit 228A can represent a unit configured to perform related operations for background transmission channel 225. For example... Figure 3A As shown in the example, the BG correlation unit 228A can perform correlation only for the background transmission channel 225 to obtain the BG correlation value 229A (which may be referred to as the BG correlation matrix 229A). Thus, the BG correlation unit 228A can perform correlation for the background transmission channel 225 individually to obtain multiple correlated background transmission channels 225, wherein the correlated background transmission channels 225 are "correlated" through the BG correlation matrix 229A. The BG correlation unit 228A can output the correlated background transmission channels 225 and the correlation matrix to the BG reordering unit 230A.
[0112] BG reordering unit 230A can represent a unit configured to reorder the relevant background transmission channels 225 based on the BG correlation matrix 229A to obtain reordered background transmission channels 231A-231D (“reordered background transmission channels 231”). As described above, BG reordering unit 230 can reorder the relevant background transmission channels 225 according to the highest correlation value between the pairs of relevant background transmission channels 225 described in the correlation matrix 229A to match the pairs of relevant background transmission channels 225.
[0113] In this way, the BG reordering unit 230A can reorder only the relevant background transmission channels 225 to obtain the reordered background transmission channel 231. The BG reordering unit 230A can output BG reordering metadata 235A to the bitstream generator 256, where the BG reordering metadata 235A can indicate how the reordered background transmission channel 231 was reordered. The BG reordering unit 230A can output the reordered background transmission channel 231 to the stereo encoder 250.
[0114] FG correlation unit 228B can operate similarly to BG correlation unit 228A, except for the foreground transmission channel 225. In this way, FG correlation unit 228B can perform correlation solely for the foreground transmission channel 225 to obtain correlation matrix 229B, and output the correlated background transmission channel and correlation matrix 229B to FG reordering unit 230B. In other words, FG correlation unit 228B can perform correlation only for the foreground transmission channel 225 to obtain correlation matrix 229B. FG correlation unit 228A can output the correlated foreground transmission channel 225 to FG reordering unit 230B.
[0115] The FG reordering unit 230B can operate similarly to the BG reordering unit 230A to reorder the relevant foreground transmission channels 225 based on the FG correlation matrix 229B, thereby obtaining the reordered foreground transmission channels 231. The FG reordering unit 230B can reorder the relevant foreground transmission channels 225 to match the relevant foreground transmission channel pairs according to the highest correlation value between the relevant foreground transmission channel pairs described in the FG correlation matrix 229B.
[0116] In this way, the FG reordering unit 230B can reorder only the relevant foreground transmission channel 225 to obtain the reordered foreground transmission channel 231. The FG reordering unit 230B can output the BG reordering metadata 235B to the bitstream generator 256, where the BG reordering metadata 235B can indicate how the reordered foreground transmission channel 231 is reordered. The FG reordering unit 230B can output the reordered foreground transmission channel 231 to the stereo encoder 250.
[0117] PAED 226A can invoke instances of stereo encoders 250A-250N (“stereo encoder 250”), which can perform psychoacoustic audio coding according to any of the stereo compression algorithms described above. Stereo encoder 250 can process two transmission channels separately to produce sub-bitstreams 233A-233N (“sub-bitstream 233”).
[0118] To compress the transmission channels, the stereo encoder 250 can perform shape and gain analysis on each reordered background and foreground transmission channel 231 to obtain the shape and gain representing the transmission channel 231. The stereo encoder 250 can also predict the first transmission channel of the pair of transmission channels 231 from the second transmission channel of the pair of transmission channels 231, and predict the gain and shape representing the first transmission channel from the gain and shape representing the second transmission channel to obtain residuals.
[0119] Before performing separate prediction on the gain, the stereo encoder 250 may first perform quantization on the gain of the second transmission channel to obtain a coarsely quantized gain and one or more finely quantized residuals. Additionally, the stereo encoder 250 may perform quantization (e.g., vector quantization) on the shape of the second transmission channel to obtain a quantized shape before performing separate prediction on the shape. The stereo encoder 250 can then use the quantization process and fine energy from the second transmission channel, as well as the quantized shape, to predict the first transmission channel from the second transmission channel, thereby predicting the quantization process and fine energy from the first transmission channel, as well as the quantized shape.
[0120] PAED 226A may also include a bitstream generator 256, which can receive sub-bitstream 233, BG reordering metadata 235A, and FG reordering metadata 235B. Bitstream generator 256 may represent a unit configured to specify sub-bitstream 233, BG reordering metadata 235A, and FG reordering metadata 235B in bitstream 231. Bitstream 231 may represent an example of bitstream 31 described above.
[0121] exist Figure 3B In the example, PAED 226B is similar to PAED 226A, except that PAED 226B includes CU 46B, wherein the combined correlation unit 228C performs correlations for all transmission channels 225 to obtain a combined correlation matrix 229C. In this respect, the combined correlation unit 228C can perform correlations for at least one background transmission channel 225 and at least one foreground transmission channel 225. The combined correlation unit 228C can output the combined correlation matrix 229C and the correlated transmission channels 225 to the combined reordering unit 230C.
[0122] Furthermore, PAED 226B differs from PAED 226A in that the combined reordering unit 230C can reorder the relevant transport channels 225 based on the combined correlation matrix 229C to obtain the reordered transport channels 231 (which may include background components and foreground audio signals or some representation thereof). The combined correlation unit 230C can determine the reordering metadata 235C, which can indicate how all transport channels 225 are reordered. The combined correlation matrix 229C can output the reordered transport channels 231 to the stereo decoder and output the reordering metadata 235C to the bitstream generator 256, both of which function as described above to generate the bitstream 231.
[0123] exist Figure 3C In the example, except for the existence of an odd number of channels that prevent reordering of one of the transmission channels 231 (i.e., Figure 3C Except for performing stereo coding on the reordered transport channel 231G in the example, PAED 226C is similar to PAED 226A. Thus, PAED 226C can invoke an instance of mono encoder 260, which can perform mono psychoacoustic audio coding for the reordered transport channel 231G, such as for... Figures 7-1 0. To be discussed in more detail.
[0124] Figure 4A and 4B It's illustrated in more detail. Figure 1 and Figure 2 The example shown is a block diagram of an example psychoacoustic audio decoding device. First refer to... Figure 4A For example, PADD 334A can represent an example of PADD 34 and / or PADD 134. PADD 334A may include bitstream extractor 338, stereo decoder 340A-340N (“stereo decoder 340”) and RU54.
[0125] Bitstream extractor 338 can represent a unit configured to parse sub-bitstream 233, reordered metadata 235 (which can refer to any of the reordered metadata 235A-235C discussed above), and ATF metadata 339 from bitstream 231. Bitstream extractor 338 can output each sub-bitstream 233 to a separate instance of stereo decoder 340. Bitstream extractor 338 can also output reordered metadata 235 to RU 54.
[0126] Each stereo decoder 340 can reconstruct the second transmission channel of the reordered transmission channel 231' pair based on the quantization gain and quantization shape described in sub-bitstream 233. Each stereo decoder 340 can then obtain a residual representing the first transmission channel of the reordered transmission channel 231' pair from sub-bitstream 233. The stereo decoder 340 can add the residual to the second transmission channel to obtain a first correlated transmission channel (e.g., correlated transmission channel 231A') from the second transmission channel (e.g., correlated transmission channel 231B'). The stereo decoder 340 can then output the reordered transmission channel 231' to RU 54.
[0127] RU 54 can reorder the reordered transport channel 231' based on reordering metadata 235. In some examples, when BG reordering metadata 235A and FG reordering metadata 235B are present in the reordering metadata 235, RU 54 can reorder the reordered background transport channel 231' separately based on BG reordering metadata 235A, and also reorder the reordered foreground transport channel 231' separately based on FG reordering metadata 235B. In other examples, multiple reordered components include a background component associated with the foreground audio signal, and RU 54 can reorder the mixture of the reordered transport channels 231' of the specified background component and the foreground audio signal based on common reordering metadata 235C. RU 54 can output the reordered transport channel 225' (which may also be referred to as ATF audio data 225') to ATF decoder 336.
[0128] ATF decoder 336 (which can perform operations similar to, if not substantially similar to, SADD 36 and / or SADD 136) can receive transport channel 225' and ATF metadata 339, and perform spatial audio decoding on transport channel 225' and the spatial components defined by ATF metadata 339 to obtain scene-based audio data 221'. Scene-based audio data 221' can represent examples of scene-based audio data 21' and / or scene-based audio data 121'.
[0129] exist Figure 4B In the example, PADD 334B is similar to PADD 334A, except that it has an odd number of channels, making it impossible to target one of the sub-bit streams 233 (i.e., Figure 4B In the example, sub-bitstream 233D) performs stereo decoding. Thus, the PADD334B can invoke an instance of the mono encoder 360, which can perform mono psychoacoustic audio encoding for sub-bitstream 233D, such as for... Figures 7-1 0. To be discussed in more detail.
[0130] Figure 5 It's illustrated in more detail. Figures 3A-3C The example shown is a block diagram of an encoder. Encoder 550 is shown as a multi-channel encoder and represents... Figure 3A and 3B The example shown is a stereo encoder 250 (where stereo encoder 250 may include only two channels, while encoder 550 has been generalized to support N channels).
[0131] like Figure 5 As shown in the example, the encoder includes gain / shape analysis units 552A-552N (“gain / shape analysis units 552”), energy quantization units 556A-556N (“energy quantization units 556”), level difference units 558A-558N (“level difference units 558”), transform units 562A-562N (“transform units 562”), and a vector quantizer 564. Each of the gain / shape analysis units 552 can be described below for the purposes of this example. Figure 7 And / or operate as described in the gain shape analysis unit in Figure 9 to perform gain shape analysis for each of the transmission channels 551 to obtain gains 553A-553N (“gain 553”) and shapes 555A-555N (“shape 555”).
[0132] The energy quantization unit 556 can be used for the following purposes: Figure 7 The encoder 550 operates as described in the energy quantizer of Figure 9 to quantize the gain 553 and thereby obtain quantized gains 557A-557N (“quantized gain 557”). The level difference unit 558 can each represent a unit configured to compare a pair of gains 553 to determine the difference between the pair. In this example, the level difference unit 558 can compare a reference gain 553A with each of the residual gains 553 to obtain gain differences 559A-559M (“gain difference 559”). The encoder 550 can specify the quantized reference gain 557A and gain difference 559 in the bitstream.
[0133] Transform unit 562 can perform subband analysis (discussed in more detail below) and apply a transform (such as KLT, which refers to the Karhunen-Loeve transform) to the subband of shape 555 to output transformed shapes 563A-563N (“transformed shape 563”). Vector quantizer 564 can perform vector quantization on the transformed shape 563 to obtain residual IDs 565A-565N (“residual ID 565”) of residual ID 565 in a specified bitstream.
[0134] The encoder 550 can also determine the combined bit allocation 560 based on the number of bits allocated to the quantization gain 557 and the gain difference 559. The combined bit allocation 560 can represent an example of the bit allocation 251 discussed in more detail above.
[0135] Figure 6 It's illustrated in more detail. Figure 4A and Figure 4B A block diagram of an example decoder. Decoder 634 is shown as a multi-channel decoder and represents... Figure 4A and Figure 4B The example shown is a stereo decoder 340 (where stereo decoder 340 may include only two channels, while decoder 634 has been generalized to support N channels).
[0136] like Figure 6 As shown in the example, decoder 634 includes level combining units 636A to 636N (“level combining units 636”), vector quantizer 638, energy dequantization units 640A to 640N (“energy dequantization units 640”), inverse transform units 642A to 642N (“transform unit 642”), and gain / shape combining units 646A to 646N (“gain / shape combining units 552”). Level combining units 636 may each represent a unit configured to combine each of the quantized reference gain 553A and gain difference 559 to determine the quantized gain 557.
[0137] The energy dequantization unit 640 can be used for the following purposes: Figure 8 And / or operate as described in the energy dequantizer of Figure 10 to dequantize the quantized gain 557 to obtain gain 553'. Encoder 550 can specify the quantization reference gain 557A and gain difference 559 in the ATF audio data.
[0138] Vector dequantizer 638 can perform vector quantization on residual ID 565 to obtain transformed shape 563'. Transform unit 562 can perform applied inverse transform (e.g., inverse KLT) and perform subband synthesis on transformed shape 563 (discussed in more detail below) to output shape 555'.
[0139] Each of the gain / shape synthesis units 552 can be targeted as follows: Figure 7 The gain / shape synthesis unit 646 operates as described in the example of Figure 9, performing gain-shape synthesis for each of gain 553' and shape 555' to obtain transmission channel 551'. The gain / shape synthesis unit 646 can output transmission channel 551' to ATF audio data.
[0140] Encoder 550 can also determine combined bit allocation 560 based on the number of bits allocated to the quantization gain 557 and gain difference 559. Combined bit allocation 560 can represent an example of bit allocation 251 discussed in more detail above.
[0141] Figure 7 The illustrations are configured to perform various aspects of the techniques described in this disclosure. Figures 3A to 3C A block diagram of an example psychoacoustic audio encoder. The audio encoder 1000A can represent an example of the PAED126, which can be configured to encode audio data for transmission via a Personal Area Network (PAN) or "PAN" (e.g., Transmission. However, the techniques of this disclosure performed by the audio encoder 1000A can be used in any context where audio data compression is required. In some examples, the audio encoder 1000A can be configured according to aptX. TM The audio codec is used to encode audio data 17, the aptX. TM Audio codecs include, for example, enhanced aptX—E-aptX, aptX Live, and aptX High Definition.
[0142] exist Figure 7 In the example, the audio encoder 1000A can be configured to encode audio data 25 using a gain-shape vector quantization encoding process, which includes encoding and decoding residual vectors using a tight mapping. During gain-shape vector quantization encoding, the audio encoder 1000A is configured to encode both the gain (e.g., energy level) and shape (e.g., residual vectors defined by transform coefficients) of a subband of the frequency domain audio data. Each subband of the frequency domain audio data represents a specific frequency range of a specific frame of the audio data 25.
[0143] Audio data 25 can be sampled at a specific sampling frequency. While any desired sampling frequency can be used, example sampling frequencies may include 48 kHz or 44.1 kHz. Each digital sample of audio data 25 can be defined by a specific input bit depth (e.g., 16 bits or 24 bits). In one example, audio encoder 1000A can be configured to operate on a single channel of audio data 21 (e.g., mono audio). In another example, audio encoder 1000A can be configured to independently encode two or more channels of audio data 25. For example, audio data 17 may include left and right channels for stereo audio. In this example, audio encoder 1000A can be configured to independently encode the left and right audio channels in a dual-mono mode. In other examples, audio encoder 1000A can be configured to encode two or more channels of audio data 25 together (e.g., in a joint stereo mode). For example, audio encoder 1000A can perform certain compression operations by predicting one channel of audio data 25 using another channel of audio data 25.
[0144] Regardless of the channel arrangement of the audio data 25, the audio encoder 1000A acquires the audio data 25 and sends it to the transformation unit 1100. The transformation unit 1100 is configured to transform the frames of the audio data 25 from the time domain to the frequency domain to produce frequency domain audio data 1112. A frame of the audio data 25 can be represented by a predetermined number of samples of the audio data. In one example, a frame of the audio data 25 can be 1024 samples wide. Different frame widths can be selected based on the frequency transformation used and the desired amount of compression. The frequency domain audio data 1112 can be represented as transformation coefficients, where the value of each transformation coefficient represents the energy of the frequency domain audio data 1112 at a specific frequency.
[0145] In one example, transform unit 1100 can be configured to transform audio data 25 into frequency domain audio data 1112 using Modified Discrete Cosine Transform (MDCT). MDCT is an "overlapping" transform based on Type IV discrete cosine transform. MDCT is considered "overlapping" when applied to data from multiple frames. That is, to perform a transform using MDCT, transform unit 1100 can include a 50% overlap window in subsequent frames of audio data. The overlap property of MDCT can be used in data compression techniques, such as audio coding, because it reduces encoding / decoding artifacts from frame boundaries. Transform unit 1100 is not limited to using MDCT and can use other frequency domain transform techniques to transform audio data 17 into frequency domain audio data 1112.
[0146] Subband filter 1102 separates frequency-domain audio data 1112 into subbands 1114. Each subband 1114 includes transform coefficients of the frequency-domain audio data 1112 within a specific frequency range. For example, subband filter 1102 can separate the frequency-domain audio data 1112 into twenty different subbands. In some examples, subband filter 1102 can be configured to separate the frequency-domain audio data 1112 into subbands 1114 with a uniform frequency range. In other examples, subband filter 1102 can be configured to separate the frequency-domain audio data 1112 into subbands 1114 with a non-uniform frequency range.
[0147] For example, subband filter 1102 can be configured to separate frequency domain audio data 1112 into subbands 1114 according to the Bark scale. Typically, Bark-scale subbands have perceptually equidistant frequency ranges. That is, Bark-scale subbands are not equal in frequency range but equal in terms of human auditory perception. Typically, subbands at lower frequencies will have fewer transform coefficients because lower frequencies are more easily perceived by the human auditory system. Thus, the frequency domain audio data 1112 in the lower frequency subbands of subband 1114 is compressed less by the audio encoder 1000A compared to the higher frequency subbands. Similarly, higher frequency subbands in subband 1114 can include more transform coefficients because higher frequencies are less perceptible to the human auditory system. Accordingly, the frequency domain audio 1112 in the data of the higher frequency subbands of subband 1114 can be compressed more by the audio encoder 1000A compared to the lower frequency subbands.
[0148] The audio encoder 1000A can be configured to process each of the sub-bands 1114 using the sub-band processing unit 1128. That is, the sub-band processing unit 1128 can be configured to process each sub-band separately. The sub-band processing unit 1128 can be configured to perform a gain shape vector quantization process with extended-range coarse-fine quantization according to the techniques of this disclosure.
[0149] The gain shape analysis unit 1104 can receive subbands 1114 as input. For each subband 1114, the gain shape analysis unit 1104 can determine the energy level 1116 of each subband 1114. That is, each subband 1114 has an associated energy level 1116. The energy level 1116 is a scalar value in decibels (dB), which represents the total energy (also called gain) in the transform coefficients of a particular subband in the subband 1114. The gain shape analysis unit 1104 can separate the energy level 1116 of one of the subbands 1114 from the transform coefficients of the subband to produce a residual vector 1118. The residual vector 1118 represents the so-called "shape" of the subband. The shape of the subband can also be referred to as the spectrum of the subband.
[0150] Vector quantizer 1108 can be configured to quantize residual vector 1118. In one example, vector quantizer 1108 can use a quantization process to quantize the residual vector to produce residual ID 1124. Instead of quantizing each sample individually (e.g., scalar quantization), vector quantizer 1108 can be configured to quantize a block of samples (e.g., a shape vector) included in the residual vector 1118. Any vector quantization technique can be used in conjunction with extended-range coarse-fine energy quantization processing.
[0151] In some examples, the audio encoder 1000A can dynamically allocate bits for encoding and decoding energy level 1116 and residual vector 1118. That is, for each of the subbands 1114, the audio encoder 1000A can determine the number of bits allocated for energy quantization (e.g., via energy quantizer 1106) and the number of bits allocated for vector quantization (e.g., via vector quantizer 1108). The total number of bits allocated for energy quantization can be referred to as the energy-allocated bits. These energy-allocated bits can then be distributed between coarse quantization and fine quantization processes.
[0152] The energy quantizer 1106 can receive the energy level 1116 of subband 1114 and quantize the energy level 1116 of subband 1114 into a coarse energy 1120 and a fine energy 1122 (which can represent one or more fine residuals of the quantization). This disclosure will describe the quantization process of a subband, but it should be understood that the energy quantizer 1106 can perform energy quantization on one or more subbands 1114 (including each subband 1114).
[0153] Typically, the energy quantizer 1106 can perform a recursive two-step quantization process. The energy quantizer 1106 can first quantize the energy level 1116 with a first number of bits for coarse quantization to produce a coarse energy 1120. The energy quantizer 1106 can use a predetermined range of energy levels for quantization (e.g., a range defined by maximum and minimum energy levels) to generate the coarse energy. The coarse energy 1120 approximates the value of energy level 1116.
[0154] Then, the energy quantizer 1106 can determine the difference between the coarse energy 1120 and the energy level 1116. This difference is sometimes referred to as the quantization error. The energy quantizer 1106 can then use a second number of bits during fine quantization to quantize the quantization error to produce a fine energy 1122. The number of bits used for fine quantization is determined by subtracting the number of bits used for coarse quantization from the total number of bits allocated to the energy. When added together, the coarse energy 1120 and the fine energy 1122 represent the total quantized value of the energy level 1116. The energy quantizer 1106 can continue to produce one or more fine energies 1122 in this manner.
[0155] The audio encoder 1000A can also be configured to encode the coarse energy 1120, fine energy 1122, and residual ID 1124 using a bitstream encoder 1110 to produce encoded audio data 31 (another way of referring to bitstream 31). The bitstream encoder 1110 can be configured to further compress the coarse energy 1120, fine energy 1122, and residual ID 1124 using one or more entropy coding processes. Entropy coding processes may include Huffman coding / decoding, arithmetic coding / decoding, context-adaptive binary arithmetic coding / decoding (CABAC), and other similar coding techniques.
[0156] In one example of this disclosure, the quantization performed by the energy quantizer 1106 is uniform quantization. That is, the step size (also called "resolution") of each quantization is equal. In some examples, the step size may be in decibels (dB). The step sizes for coarse quantization and fine quantization may be determined from a predetermined range of energy values used for quantization and the number of bits allocated for quantization, respectively. In one example, the energy quantizer 1106 performs uniform quantization on both coarse quantization (e.g., producing coarse energy 1120) and fine quantization (e.g., producing fine energy 1122).
[0157] Performing a two-step uniform quantization process is equivalent to performing a single uniform quantization process. However, by splitting the uniform quantization into two parts, the bits allocated to coarse quantization and fine quantization can be controlled independently. This allows for greater flexibility in bit allocation across energy and vector quantization and can improve compression efficiency. Consider an M-level uniform quantizer, where M defines the number of levels (e.g., in dB), the number of energy levels that can be divided into. M can be determined by the number of bits allocated for quantization. For example, the energy quantizer 1106 can use M1 levels for coarse quantization and M2 levels for fine quantization. This is equivalent to a single uniform quantizer using M1*M2 levels.
[0158] Figure 8 It's illustrated in more detail. Figure 4A and Figure 4B A block diagram illustrating the implementation of a psychoacoustic audio decoder. Audio decoder 1002A can represent an example of decoder 510, which can be configured to output data via a PAN (e.g., ...). The received audio data is decoded. However, the techniques of this disclosure performed by the audio decoder 1002A can be used in any context where audio data compression is required. In some examples, the audio decoder 1002A can be configured according to aptX. TM The audio codec is used to decode audio data 21, by aptX. TMAudio codecs include, for example, enhanced aptX—E-aptX, aptX Live, and aptX High Definition. However, the techniques disclosed herein can be used in any audio codec configured to perform quantization of audio data. Audio decoder 1002A can be configured to perform various aspects of the quantization process using tight mapping according to the techniques disclosed herein.
[0159] Typically, the audio decoder 1002A can operate reciprocally with respect to the audio encoder 1000A. This allows the same procedures used in encoders for quality / bitrate scalable cooperative PVQ to be applied in the audio decoder 1002A. Decoding is based on the same principles, but in reverse to the operations performed in the decoder, enabling the reconstruction of audio data from the encoded bitstream received from the autoencoder. Each quantizer has an associated dequantizer corresponding section. For example, as... Figure 8 As shown, the inverse transform unit 1100', inverse subband filter 1102', gain shape synthesis unit 1104', energy dequantizer 1106', vector dequantizer 1108', and bitstream decoder 1110' can be configured to perform actions targeting... Figure 7 The inverse operation of the transformation unit 1100, subband filter 1102, gain shape analysis unit 1104, energy quantizer 1106, vector quantizer 1108 and bitstream encoder 1110.
[0160] Specifically, the gain-shape synthesis unit 1104' reconstructs the frequency domain audio data, which has a reconstructed residual vector and a reconstructed energy level. The inverse subband filter 1102' and the inverse transform unit 1100' output the reconstructed audio data 25'. In an example where the encoding is lossless, the reconstructed audio data 25' perfectly matches the audio data 25. In an example where the encoding is lossy, the reconstructed audio data 25' may not perfectly match the audio data 25.
[0161] Figure 9A and Figure 9B It's illustrated in more detail. Figures 3A-3C A block diagram of an additional example of a psychoacoustic audio encoder is shown in the example. First, refer to... Figure 9A For example, the audio encoder 1000B can be configured to encode audio data for use via a PAN (e.g., Transmission. However, similarly, the techniques of this disclosure performed by the audio encoder 1000B can be used in any context where audio data compression is required. In some examples, the audio encoder 1000B can be configured according to aptX. TM The audio codec is used to encode audio data 25, the aptX. TMAudio codecs include, for example, enhanced aptX—E-aptX, aptX Live, and aptX High Definition. However, the techniques disclosed herein can be used in any audio codec. As will be explained in more detail below, the audio encoder 1000B can be configured to perform various aspects of perceptual audio encoding and decoding according to the various aspects of the techniques described in this disclosure.
[0162] exist Figure 9A In the example, the audio encoder 1000B can be configured to encode audio data 25 using a gain-shape vector quantization encoding process. During gain-shape vector quantization encoding, the audio encoder 1000B is configured to encode both the gain (e.g., energy level) and shape (e.g., a residual vector defined by the transform coefficients) of a subband of the frequency domain audio data. Each subband of the frequency domain audio data represents a specific frequency range of a specific frame of the audio data 25. Generally, throughout this disclosure, the term "subband" refers to a frequency range, frequency band, etc.
[0163] The audio encoder 1000B invokes the transformation unit 1100 to process the audio data 25. The transformation unit 1100 is configured to process the audio data 25 by applying a transformation to at least partially the frames of the audio data 25, thereby transforming the audio data 25 from the time domain to the frequency domain to produce frequency domain audio data 1112.
[0164] A frame of audio data 25 can be represented by a predetermined number of samples of the audio data. In one example, a frame of audio data 25 can be 1024 samples wide. Different frame widths can be chosen based on the frequency transform used and the desired amount of compression. Frequency domain audio data 1112 can be represented as transform coefficients, where the value of each transform coefficient represents the energy of the frequency domain audio data 1112 at a specific frequency.
[0165] In one example, transform unit 1100 can be configured to transform audio data 25 into frequency domain audio data 1112 using Modified Discrete Cosine Transform (MDCT). MDCT is an "overlapping" transform based on Type IV discrete cosine transform. MDCT is considered "overlapping" when applied to data from multiple frames. That is, to perform a transform using MDCT, transform unit 1100 can include a 50% overlap window in subsequent frames of audio data. The overlap property of MDCT can be used in data compression techniques, such as audio coding, because it can reduce encoding / decoding artifacts from frame boundaries. Transform unit 1100 is not limited to using MDCT and can use other frequency domain transform techniques to transform audio data 25 into frequency domain audio data 1112.
[0166] Subband filter 1102 separates frequency-domain audio data 1112 into subbands 1114. Each subband 1114 includes transform coefficients of the frequency-domain audio data 1112 within a specific frequency range. For example, subband filter 1102 can separate the frequency-domain audio data 1112 into twenty different subbands. In some examples, subband filter 1102 can be configured to separate the frequency-domain audio data 1112 into subbands 1114 with a uniform frequency range. In other examples, subband filter 1102 can be configured to separate the frequency-domain audio data 1112 into subbands 1114 with a non-uniform frequency range.
[0167] For example, subband filter 1102 can be configured to separate frequency domain audio data 1112 into subbands 1114 according to the Bark scale. Typically, Bark-scale subbands have perceptually equidistant frequency ranges. That is, Bark-scale subbands are not equal in frequency range, but equal in terms of human auditory perception. Typically, subbands at lower frequencies will have fewer transform coefficients because lower frequencies are more easily perceived by the human auditory system.
[0168] Thus, the frequency domain audio data 1112 in the lower frequency subband of subband 1114 is compressed less by the audio encoder 1000B compared to the higher frequency subband. Similarly, the higher frequency subband of subband 1114 can include more transform coefficients because higher frequencies are more difficult for the human auditory system to perceive. Therefore, the frequency domain audio 1112 in the data of the higher frequency subband of subband 1114 can be compressed more by the audio encoder 1000B compared to the lower frequency subband.
[0169] The audio encoder 1000B can be configured to process each of the sub-bands 1114 using the sub-band processing unit 1128. That is, the sub-band processing unit 1128 can be configured to process each sub-band separately. The sub-band processing unit 1128 can be configured to perform gain shape vector quantization processing.
[0170] The gain shape analysis unit 1104 can receive subbands 1114 as input. For each subband 1114, the gain shape analysis unit 1104 can determine the energy level 1116 of each subband 1114. That is, each subband 1114 has an associated energy level 1116. The energy level 1116 is a scalar value in decibels (dB), which represents the total energy (also called gain) in the transform coefficients of a particular subband 1114. The gain shape analysis unit 1104 can separate the energy level 1116 of one of the subbands 1114 from the transform coefficients of the subband to produce a residual vector 1118. The residual vector 1118 represents the so-called "shape" of the subband. The shape of the subband can also be referred to as the spectrum of the subband.
[0171] Vector quantizer 1108 can be configured to quantize residual vector 1118. In one example, vector quantizer 1108 can use a quantization process to quantize the residual vector to produce residual ID 1124. Instead of quantizing each sample individually (e.g., scalar quantization), vector quantizer 1108 can be configured to quantize a block of samples (e.g., a shape vector) included in the residual vector 1118.
[0172] In some examples, the audio encoder 1000B can dynamically allocate bits for encoding energy level 1116 and residual vector 1118. That is, for each of the subbands 1114, the audio encoder 1000B can determine the number of bits allocated for energy quantization (e.g., via energy quantizer 1106) and the number of bits allocated for vector quantization (e.g., via vector quantizer 1108). The total number of bits allocated for energy quantization can be referred to as the energy-allocated bits. These energy-allocated bits can then be distributed between coarse quantization and fine quantization processes.
[0173] The energy quantizer 1106 can receive the energy level 1116 of the subband 1114 and quantize the energy level 1116 of the subband 1114 into a coarse energy 1120 and a fine energy 1122. This disclosure will describe the quantization process of a subband, but it should be understood that the energy quantizer 1106 can perform energy quantization on one or more subbands 1114 (including each subband 1114).
[0174] like Figure 9A As shown in the example, the energy quantizer 1106 may include a prediction / difference (“P / D”) unit 1130, a coarse quantization (“CQ”) unit 1132, a summation unit 1134, and a fine quantization (“FQ”) unit 1136. The P / D unit 1130 can predict or otherwise identify the difference 1116 between the energy levels of one sub-band 1114 and another sub-band 1114 within the same frame of audio data (which may refer to spatial prediction in the frequency domain), or the difference 1116 between the same (or possibly different) energy levels of one sub-band 1114 from different frames (which may be referred to as temporal prediction). The P / D unit 1130 can analyze the energy levels 1116 in this way to obtain a predicted energy level 1131 (“PEL 1131”) for each sub-band 1114. The P / D unit 1130 can output the predicted energy level 1131 to the coarse quantization unit 1132.
[0175] Coarse quantization unit 1132 can be represented as a unit configured to perform coarse quantization on a predicted energy level 1131 to obtain a coarse energy 1120. Coarse quantization unit 1132 can output the coarse energy 1120 to bitstream encoder 1110 and summing unit 1134. Summing unit 1134 can be represented as a unit configured to obtain the difference between coarse quantization unit 1134 and predicted energy level 1131. Summing unit 1134 can output the difference as an error 1135 (which may also be referred to as "residual 1135") to fine quantization unit 1135.
[0176] Fine quantization unit 1132 can represent a unit configured to perform fine quantization on error 1135. Fine quantization can be considered "fine" compared to coarse quantization performed by coarse quantization unit 1132. That is, fine quantization unit 1132 can quantize based on a step size with higher resolution than the step size used when performing coarse quantization, thereby quantizing error 1135. As a result of performing fine quantization on error 1135, fine quantization unit 1136 can obtain fine energy 1122 for each of sub-bands 1122. Fine quantization unit 1136 can output the fine energy 1122 to bitstream encoder 1110.
[0177] Typically, the energy quantizer 1106 can perform a multi-step quantization process. The energy quantizer 1106 can first quantize the energy level 1116 with a first number of bits for coarse quantization to produce a coarse energy 1120. The energy quantizer 1106 can use a predetermined range of energy levels for quantization (e.g., a range defined by the maximum and minimum energy levels) to generate the coarse energy. The coarse energy 1120 approximates the value of energy level 1116.
[0178] The energy quantizer 1106 then determines the difference between the coarse energy 1120 and the energy level 1116. This difference is sometimes referred to as the quantization error (or residual). The energy quantizer 1106 can then use a second number of bits during the fine quantization process to quantize the quantization error, producing the fine energy 1122. The number of bits used for fine quantization is determined by subtracting the number of bits used for coarse quantization from the total number of bits allocated to the energy. When added together, the coarse energy 1120 and the fine energy 1122 represent the total quantized value of the energy level 1116.
[0179] The audio encoder 1000B can also be configured to encode the coarse energy 1120, fine energy 1122, and residual ID 1124 using the bitstream encoder 1110 to produce encoded audio data 21. The bitstream encoder 1110 can be configured to further compress the coarse energy 1120, fine energy 1122, and residual ID 1124 using one or more of the entropy encoding processes described above.
[0180] According to aspects of this disclosure, the energy quantizer 1106 (and / or its components, such as the fine quantization unit 1136) can implement a hierarchical rate control mechanism to provide a high degree of scalability and achieve seamless or substantially seamless real-time streaming. For example, the fine quantization unit 1136 can implement a hierarchical fine quantization scheme according to aspects of this disclosure. In some examples, the fine quantization unit 1136 invokes a multiplexer (or “MUX”) 1137 to implement the selection operation of the hierarchical rate control.
[0181] The term "coarse quantization" refers to the combined operation of the two coarse-fine quantization processes described above. According to various aspects of this disclosure, the fine quantization unit 1136 can perform one or more additional iterations of fine quantization on the error 1135 received from the summing unit 1134. The fine quantization unit 1136 can use a multiplexer 1137 to switch between and traverse various (more)fine energy levels.
[0182] Hierarchical rate control can refer to a tree-based fine quantization structure or a cascaded fine quantization structure. When considered as a tree-based structure, the existing two-step quantization operation forms the root node of the tree, and this root node is described as having a resolution depth of one (1). According to the techniques of this disclosure, depending on the availability of bits for further fine quantization, multiplexer 1137 can select one or more additional levels of fine quantization. Any of these subsequent fine quantization levels selected by multiplexer 1137 represents a resolution depth of two (2), three (3) for a tree-based structure representing the multi-level fine quantization technique of this disclosure.
[0183] The fine quantization unit 1136 can provide improved scalability and control for seamless real-time streaming scenarios in wireless PANs. For example, the fine quantization unit 1136 can replicate hierarchical fine quantization schemes and quantization multiplexing trees at higher levels, seeding at coarse quantization points in a more general decision tree. Furthermore, the fine quantization unit 1136 enables the audio encoder 1000B to achieve seamless or substantially seamless real-time compression and streaming navigation. For example, the fine quantization unit 1136 can perform multi-root hierarchical decision structures for multi-level fine quantization, allowing the energy quantizer 1106 to utilize the total available bits to perform a potential number of iterations of fine quantization.
[0184] The fine quantization unit 1136 can implement the hierarchical rate control process in various ways. The fine quantization unit 1136 can invoke the multiplexer 1137 on a per-subband basis to independently multiplex (and thereby select the appropriate tree-based quantization scheme) the error 1135 information for each of the subbands 1114. That is, in these examples, the fine quantization unit 1136 performs a multiplexing-based hierarchical quantization mechanism selection for each corresponding subband 1114 independently of the quantization mechanism selection for any other subband in the subband 1114. In these examples, the fine quantization unit 1136 quantizes each of the subbands 1114 only according to the target bit rate specified for the corresponding subband 1114. In these examples, the audio encoder 1000B can signal details of the specific hierarchical quantization scheme for each subband 1114 as part of the encoded audio data 21.
[0185] In other examples, the fine quantization unit 1136 may invoke the multiplexer 1137 only once, thereby selecting a single multiplex-based quantization scheme for the error information 1135 belonging to all sub-bands 1114. That is, in these examples, the fine quantization unit 1136 quantizes the error information 1135 belonging to all sub-bands 1114 according to the same target bit rate, which is selected in a single time and uniformly defined for all sub-bands 1114. In these examples, the audio encoder 1000B may signal, as part of the encoded audio data 21, the details of the single layered quantization scheme applied to all sub-bands 1114.
[0186] Next reference Figure 9B For example, the audio encoder 1000C can represent Figure 1 and Figure 2 The example shown is another example of the psychoacoustic audio encoding devices 26 and / or 126. The audio encoder 1000C is similar to... Figure 9A The audio encoder 1000B shown in the example, except for the audio encoder 1000C, includes a general analysis unit 1148, a quantization controller unit 1150, a general quantizer 1156, and a cognitive / perceptual / auditory / psychoacoustic (CPHP) quantizer 1160 that can perform gain synthesis analysis or any other type of analysis on the output level 1149 and the residual 1151.
[0187] The general analysis unit 1148 can receive sub-band 1114 and perform any type of analysis to generate a level 1149 and a residual 1151. The general analysis unit 1148 can output level 1149 to the quantization controller unit 1150 and residual 1151 to the CPHP quantizer 1160.
[0188] The quantization controller unit 1150 can receive level 1149. For example... Figure 9B As shown in the example, the quantization controller unit 1150 may include a hierarchical specification unit 1152 and a specification control (SC) manager unit 1154. In response to reception level 1149, the quantization controller unit 1150 may invoke the hierarchical specification unit 1152, which may perform top-down or bottom-up hierarchical specification. Figure 11 This is an illustration of an example of top-down quantization. Figure 12 This is an illustration of an example of bottom-up quantization. That is, the hierarchical canonical unit 1152 can switch back and forth between coarse quantization and fine quantization on a frame-by-frame basis to enable a weighting mechanism that can make any given quantization coarser or finer.
[0189] The transition from a coarse to a finer state can occur by weighting the previous quantization error. Alternatively, quantization can occur such that adjacent quantization points are grouped together into a single quantization point (moving from a fine to a coarse state). Such an implementation can use sequential data structures, such as linked lists, or richer structures, such as trees or graphs. Thus, the hierarchical specification unit 1152 can determine whether to switch from fine to coarse quantization or from coarse to fine quantization, thereby providing the hierarchical space 1153 (which is the set of quantization points for the current frame) to the SC manager unit 1154. The hierarchical specification unit 1152 can determine whether to switch between finer and coarser quantization based on any information used to perform the fine or coarse quantization specified above (e.g., temporal or spatial priority information).
[0190] The SC manager unit 1154 can receive the layer space 1153 and generate specification metadata 1155, and pass the instruction 1159 of the layer space 1153 together with the specification metadata 1155 to the bitstream encoder 1110. The SC manager unit 1154 can also output the layer specification 1159 to the quantizer 1156, which performs quantization on level 1149 according to the layer space 1159 to obtain quantization level 1157. The quantizer 1156 can output the quantization level 1157 to the bitstream encoder 1110, which can operate as described above to form encoded audio data 31.
[0191] CPHP quantizer 1160 can perform one or more of cognitive, perceptual, auditory, and psychoacoustic coding on residual 1151 to obtain residual ID 1161. CPHP quantizer 1160 can output residual ID 1161 to bitstream encoder 1110, which can operate as described above to form encoded audio data 31.
[0192] Figure 10A and 10B It's illustrated in more detail. Figure 4A and 4B A block diagram of an additional example of a psychoacoustic audio decoder. Figure 10A In the example, audio decoder 1002B represents Figure 3A Another example of the AptX decoder 510 shown in the example. The audio decoder 1002B includes an extraction unit 1232, a subband reconstruction unit 1234, and a reconstruction unit 1236. The extraction unit 1232 may represent a unit configured to extract coarse energy 1120, fine energy 1122, and residual ID 1124 from encoded audio data 31. The extraction unit 1232 may extract one or more of the coarse energy 1120, fine energy 1122, and residual ID 1124 based on energy bit allocation 1203. The extraction unit 1232 may output the coarse energy 1120, fine energy 1122, and residual ID 1124 to the subband reconstruction unit 1234.
[0193] Subband reconstruction unit 1234 can be represented as being configured to be with Figure 9A The example illustrates a sub-band processing unit 1128 of the audio encoder 1000B that operates in a reciprocal manner. In other words, the sub-band reconstruction unit 1234 can reconstruct a sub-band from coarse energy 1120, fine energy 1122, and residual ID 1124. The sub-band reconstruction unit 1234 may include an energy dequantizer 1238, a vector dequantizer 1240, and a sub-band synthesizer 1242.
[0194] Energy dequantizer 1238 can be represented as being configured to... Figure 9A The energy quantizer 1106 shown performs dequantization in a quantization reciprocal manner. The energy dequantizer 1238 can perform dequantization on coarse and fine energies 1122 to obtain a predicted / differential energy level. The energy dequantizer 1238 can perform inverse prediction or differential calculation to obtain energy level 1116. The energy dequantizer 1238 can output energy level 1116 to the subband synthesizer 1242.
[0195] If the encoded audio data 31 includes a syntax element that indicates the fine energy 1122 is to be dequantized in layers, then the energy dequantizer 1238 can dequantize the fine energy 1122 in layers. In some examples, the encoded audio data 31 may include a syntax element that indicates whether the fine energy 1122 is formed using the same layered quantization structure across all sub-bands 1114, or whether a separate layered quantization structure is determined for each sub-band 1114. Based on the value of the syntax element, the energy dequantizer 1238 may apply the same layered dequantization structure to all sub-bands 1114 represented by the fine energy 1122, or it may update the layered dequantization structure based on each sub-band when dequantizing the fine energy 1122.
[0196] Vector dequantizer 1240 can represent a unit configured to perform vector dequantization in a manner reciprocal to vector quantization performed by vector quantizer 1108. Vector dequantizer 1240 can perform vector dequantization on residual ID 1124 to obtain residual vector 1118. Vector dequantizer 1240 can output residual vector 1118 to subband synthesizer 1242.
[0197] Subband synthesizer 1242 can represent a unit configured to operate reciprocally with gain shape analysis unit 1104. Thus, subband synthesizer 1242 can perform inverse gain shape analysis for energy level 1116 and residual vector 1118 to obtain subband 1114. Subband synthesizer 1242 can output subband 1114 to reconstruction unit 1236.
[0198] Reconstruction unit 1236 can represent a unit configured to reconstruct audio data 25' based on subband 1114. In other words, reconstruction unit 1236 can perform inverse subband filtering in a manner reciprocal with the subband filtering applied by subband filter 1102 to obtain frequency domain audio data 1112. Reconstruction unit 1236 can then perform inverse transform in a manner reciprocal with the transform applied by transform unit 1100 to obtain audio data 25'.
[0199] Next reference Figure 10B For example, the audio decoder 1002C can represent Figure 1 and / or Figure 2 The example shown is one example of the psychoacoustic audio decoding devices 34 and / or 134. Furthermore, the audio decoder 1002C may be similar to the audio decoder 1002B, except that the audio decoder 1002C may include an abstract control manager 1250, a hierarchical abstraction unit 1252, a dequantizer 1254, and a CPHP dequantizer 1256.
[0200] Abstract control manager 1250 and hierarchical abstraction unit 1252 can form dequantizer controller 1249, which controls the operation of dequantizer 1254 and is interoperable with quantizer controller 1150. Thus, abstract control manager 1250 can interoperate with SC manager unit 1154, receiving metadata 1155 and hierarchical specification 1159. Abstract control manager 1250 processes metadata 1155 and hierarchical specification 1159 to obtain hierarchical space 1153, which is then output to hierarchical abstraction unit 1252. Hierarchical abstraction unit 1252 is interoperable with hierarchical specification unit 1152, thereby processing hierarchical space 1153 to output indication 1159 of hierarchical space 1153 to dequantizer 1254.
[0201] Dequantizer 1254 can interoperate with quantizer 1156, wherein dequantizer 1254 can use the instruction 1159 of layer space 1153 to dequantize quantization level 1157 to obtain dequantization level 1149. Dequantizer 1254 can output dequantization level 1149 to subband synthesizer 1242.
[0202] Extraction unit 1232 can output residual ID 1161 to CPHP dequantizer 1256, which is interoperable with CPHP quantizer 1160. CPHP dequantizer 1256 can process residual ID 1161 to dequantize residual ID 1161 and obtain residual 1161. CPHP dequantizer 1256 can output the residual to subband synthesizer 1242, which can process residual 1151 and dequantization stage 1254 to output subband 1114. Reconstruction unit 1236 can operate as described above to convert subband 1114 into audio data 25' by applying an inverse subband filter to subband 1114 and then applying an inverse transform to the output of the inverse subband filter.
[0203] Figure 13 It's a diagram. Figure 2 The example shown is a block diagram of an example component of the source device. Figure 13 In the example, source device 112 includes a processor 412, a graphics processing unit (GPU) 414, system memory 416, a display processor 418, one or more integrated speakers 140, a display 103, a user interface 420, an antenna 421, and a transceiver module 422. In the example where source device 112 is a mobile device, the display processor 418 is a mobile display processor (MDP). In some examples, such as in the example where source device 112 is a mobile device, processor 412, GPU 414, and display processor 418 may be formed as an integrated circuit (IC).
[0204] For example, an IC can be considered a processing chip within a chip package and can be a system-on-a-chip (SoC). In some examples, two of the processor 412, GPU 414, and display processor 418 may be housed together in the same IC, while another may be housed in a different integrated circuit (i.e., a different chip package), or all three may be housed in different ICs or on the same IC. However, in the example where the source device 12 is a mobile device, the processor 412, GPU 414, and display processor 418 may all be housed in different integrated circuits.
[0205] Examples of processor 412, GPU 414, and display processor 418 include (but are not limited to) one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Processor 412 may be the central processing unit (CPU) of source device 12. In some examples, GPU 414 may be dedicated hardware including integrated and / or discrete logic circuits that provide GPU 414 with substantial parallel processing capabilities suitable for graphics processing. In some cases, GPU 414 may also include general-purpose processing capabilities and may be referred to as a general-purpose GPU (GPGPU) when implementing general-purpose processing tasks (i.e., non-graphics-related tasks). Display processor 418 may also be dedicated integrated circuit hardware designed to retrieve image content from system memory 416, assemble the image content into image frames, and output the image frames to display 103.
[0206] Processor 412 can execute various types of applications 20. Examples of applications 20 include web browsers, email applications, spreadsheets, video games, other applications that generate visual objects for display, or any of the application types listed above in more detail. System memory 416 can store instructions for executing applications 20. Executing one of the applications 20 on processor 412 causes processor 412 to generate graphical data of image content to be displayed and audio data 21 to be played (possibly via integrated speaker 105). Processor 412 can transfer the graphical data of the image content to GPU 414 for further processing based on the instructions or commands transferred from processor 412 to GPU 414.
[0207] The processor 412 can communicate with the GPU 414 according to a specific application processing interface (API). Examples of such APIs include... of API, Khronos Group or OpenGL and OpenCL TMHowever, aspects of this disclosure are not limited to DirectX, OpenGL, or OpenCL APIs and can be extended to other types of APIs. Furthermore, the techniques described in this disclosure do not require API compatibility to function, and the processor 412 and GPU 414 can communicate using any technology.
[0208] System memory 416 may be the memory of source device 12. System memory 416 may include one or more computer-readable storage media. Examples of system memory 416 include, but are not limited to, random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other media that can be used to carry or store desired program code in the form of instructions and / or data structures and that can be accessed by a computer or processor.
[0209] In some examples, system memory 416 may include instructions that cause processor 412, GPU 414, and / or display processor 418 to perform the functions assigned to processor 412, GPU 414, and / or display processor 418 in this disclosure. Therefore, system memory 416 may be a computer-readable storage medium on which instructions are stored, which, when executed, cause one or more processors (e.g., processor 412, GPU 414, and / or display processor 418) to perform various functions.
[0210] System memory 416 may include a non-transitory storage medium. The term "non-transitory" means that the storage medium is not included in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that system memory 416 is immovable or that its contents are static. As an example, system memory 416 may be removed from source device 12 and moved to another device. As another example, a memory substantially similar to system memory 416 may be inserted into source device 12. In some examples, the non-transitory storage medium may store data that may change over time (e.g., in RAM).
[0211] User interface 420 may represent one or more hardware or virtual (meaning a combination of hardware and software) user interfaces through which a user can connect to source device 12. User interface 420 may include physical buttons, switches, triggers, lights, or virtual versions thereof. User interface 420 may also include physical or virtual keyboards, such as the touch interface of a touchscreen, haptic feedback, etc.
[0212] Processor 412 may include one or more hardware units (including so-called "processing cores") configured to perform all or some of the operations discussed above for one or more of the mixing unit 120, audio encoder 122, wireless connection manager 128, and wireless communication unit 130. Antenna 421 and transceiver module 422 may represent units configured to establish and maintain a wireless connection between source device 12 and sink device 114. Antenna 421 and transceiver module 422 may represent one or more receivers and / or one or more transmitters capable of wireless communication according to one or more wireless communication protocols. That is, transceiver module 422 may represent a separate transmitter, a separate receiver, both a separate transmitter and a separate receiver, or a combination of transmitters and receivers. Antenna 421 and transceiver 422 may be configured to receive encoded audio data encoded according to the techniques of this disclosure. Similarly, antenna 421 and transceiver 422 may be configured to transmit encoded audio data encoded according to the techniques of this disclosure. The transceiver module 422 can perform all or part of the operations of one or more of the wireless connection manager 128 and the wireless communication unit 130.
[0213] Figure 14 It's a diagram. Figure 2 The block diagram of exemplary components of the sink device shown in the example. Although the sink device 114 may include components related to those described above. Figure 13 The examples discussed in more detail are similar to the components of source device 112, but in some cases, sink device 14 may include only a subset of the components discussed above for source device 112.
[0214] exist Figure 14 In the example, sink device 114 includes one or more speakers 802, a processor 812, system memory 816, a user interface 820, an antenna 821, and a transceiver module 822. The processor 812 may be similar to or substantially similar to processor 412. In some cases, the processor 812 may differ from processor 412 in terms of total processing power, or may be customized for low power consumption. The system memory 816 may be similar to or substantially similar to system memory 416. The speaker 140, user interface 820, antenna 821, and transceiver module 822 may be similar to or substantially similar to the corresponding speaker 440, user interface 420, and transceiver module 422. While sink device 114 may optionally include a display 800, the display 800 may represent a low-power, low-resolution (possibly monochrome LED) display through which limited information that can be directly driven by the processor 812 is transmitted.
[0215] Processor 812 may include one or more hardware units (including so-called "processing cores") configured to perform all or some of the operations discussed above for one or more of the wireless connection manager 150, wireless communication unit 152, and audio decoder 132. Antenna 821 and transceiver module 822 may represent units configured to establish and maintain a wireless connection between source device 112 and sink device 114. Antenna 821 and transceiver module 822 may represent one or more receivers and one or more transmitters capable of wireless communication according to one or more wireless communication protocols. Antenna 821 and transceiver 822 may be configured to receive encoded audio data encoded according to the techniques of this disclosure. Similarly, antenna 821 and transceiver 822 may be configured to transmit encoded audio data encoded according to the techniques of this disclosure. Transceiver module 822 may perform all or some of the operations of one or more of the wireless connection manager 150 and wireless communication unit 152.
[0216] Figure 15 It's a diagram. Figure 1 The example shown is a flowchart illustrating exemplary operation of an audio encoder performing various aspects of the techniques described in this disclosure. Audio encoder 22 may invoke spatial audio encoding device 24, which may perform spatial audio encoding on scene-based audio data 21 to obtain multiple background components, multiple foreground audio signals, and corresponding multiple spatial components as ATF audio data 25 (1300). Spatial audio encoding device 24 may output ATF audio data 25 to psychoacoustic audio encoding device 26.
[0217] The psychoacoustic audio encoding device 26 can receive background and foreground audio signals. For example... Figure 1 As illustrated in the example, the psychoacoustic audio coding device 26 may include a correlation unit (CU) 46, which can perform the aforementioned correlation with two or more of a plurality of background components and a plurality of foreground audio signals to obtain a plurality of correlated components (1302). As described above, the CU 46 may perform correlation with the background components and the foreground audio signals separately. In other examples, as described above, the CU 46 may perform correlation with both the background components and the foreground audio signals, wherein at least one background component and at least one foreground audio signal undergo correlation.
[0218] CU 46 can obtain correlation values as a result of performing correlation on two or more background components and foreground audio signals. CU 46 can reorder the background components and foreground audio signals based on the correlation values. CU 46 can output reordering metadata indicating how the transport channel is reordered to spatial audio coding device 24, which can specify the reordering metadata in metadata including spatial components. Although described as specifying reordering metadata in metadata including spatial components, psychoacoustic audio coding device 26 can specify the reordering metadata in bitstream 31.
[0219] In any case, CU 46 can output a reordered transmission channel (which can specify multiple related components), so that psychoacoustic audio coding device 26 can perform psychoacoustic audio coding on the multiple related components to obtain the encoded components (1304). Psychoacoustic audio coding device 26 can specify multiple encoded components in bitstream 31 (1306).
[0220] Figure 16 It's a diagram. Figure 1 The example shown is a flowchart illustrating exemplary operation of the audio decoder when performing various aspects of the techniques described in this disclosure. As described above, the audio decoder 32 is interoperable with the audio encoder 22. Thus, the audio decoder 32 can obtain multiple encoded related components (1400) from the bitstream 31. The audio decoder 32 can invoke the psychoacoustic audio decoding device 34, which can perform psychoacoustic audio decoding on one or more of the multiple encoded related components to obtain multiple related components (1401).
[0221] The psychoacoustic audio decoding device 34 can obtain reordering metadata from the bitstream 31, which can represent an indication of how one or more of a plurality of related components are reordered in the bitstream 31 (1402). Figure 1 As illustrated in the example, the psychoacoustic audio decoding device 34 may include a reordering unit (RU) 54, which represents a unit (1404) configured to reorder multiple related components based on reordering metadata to obtain multiple reordered components. The psychoacoustic audio decoding device 34 may reconstruct ATF audio data 25' based on the multiple reordered components (1406). The spatial audio decoding device 36 may then reconstruct scene-based audio data 21' based on the ATF audio data 25'.
[0222] The foregoing aspects of the technology can be implemented in accordance with the following provisions.
[0223] Clause 1F. An apparatus configured to encode scene-based audio data, the apparatus comprising: a memory configured to store the scene-based audio data; and one or more processors configured to: perform spatial audio coding on the scene-based audio data to obtain a plurality of background components, a plurality of foreground audio signals, and a plurality of corresponding spatial components of a sound field represented by the scene-based audio data, each of the plurality of spatial components defining spatial characteristics of a corresponding foreground audio signal in the plurality of foreground audio signals; perform correlation with two or more of the plurality of background components and the plurality of foreground audio signals to obtain a plurality of correlated components; perform psychoacoustic audio coding / decoding with one or more of the plurality of correlated components to obtain encoded components; and specify the encoded components in a bitstream.
[0224] Clause 2F. A device according to Clause 1F, wherein the one or more processors are configured to perform psychoacoustic audio coding according to the AptX compression algorithm for at least one pair of correlated components among the plurality of correlated components.
[0225] Clause 3F. Any combination of the devices pursuant to Clauses 1F and 2F, wherein the one or more processors are configured to perform psychoacoustic audio coding for at least one pair of the plurality of related components to obtain the encoded components.
[0226] Clause 4F. A device according to any combination of Clauses 1F-3F, wherein the one or more processors are configured to: perform correlation separately for the plurality of background components to obtain a plurality of correlated background components; and perform psychoacoustic audio coding for at least one pair of the plurality of background components.
[0227] Clause 5F. An apparatus according to any combination of Clauses 1F-4F, wherein the one or more processors are configured to: perform correlation individually on the plurality of foreground audio signals to obtain a plurality of correlated foreground audio signals of the plurality of correlated components; and perform psychoacoustic audio encoding on at least one pair of the plurality of correlated foreground audio signals.
[0228] Clause 6F. A device according to any combination of Clauses 1F-5F, wherein the one or more processors are configured to perform correlation with at least one of the plurality of background components and at least one of the plurality of foreground audio signals to obtain at least one pair of the plurality of correlated components.
[0229] Clause 7F. A device pursuant to any combination of Clauses 1F-6F, wherein the one or more processors are further configured to: reorder one or more of the plurality of background components and the plurality of foreground audio signals in the bitstream based on the correlation; and specify in the bitstream an indication of how the one or more of the plurality of background components of the plurality of foreground audio signals are reordered in the bitstream.
[0230] Clause 8F. A device according to any combination of Clauses 1F to 7F, wherein the one or more processors are configured to perform a linear reversible transformation on the scene-based audio data to obtain the plurality of foreground audio signals and the corresponding plurality of spatial components.
[0231] Clause 9F. Any device according to any combination of Clauses 1F-8F, wherein the scene-based audio data includes higher-order high-fidelity stereo coefficients corresponding to orders greater than zero.
[0232] Clause 10F. Any device according to a combination of Clauses 1F-9F, wherein the scene-based audio data includes audio data defined in the spherical harmonic domain.
[0233] Clause 11F. Any device according to any combination of Clauses 1F-10F, wherein each of the plurality of foreground audio signals includes a foreground audio signal defined in the spherical harmonic domain, and wherein each of the corresponding plurality of spatial components includes a spatial component defined in the spherical harmonic domain.
[0234] Clause 12F. A method for encoding scene-based audio data, the method comprising: performing spatial audio coding on the scene-based audio data to obtain a plurality of background components, a plurality of foreground audio signals, and a plurality of corresponding spatial components of a sound field represented by the scene-based audio data, each of the plurality of spatial components defining spatial characteristics of a corresponding foreground audio signal in the plurality of foreground audio signals; correlating the plurality of background components and one or more of the plurality of foreground audio signals to obtain a plurality of correlated components; performing psychoacoustic audio coding on one or more of the plurality of correlated components to obtain an encoded component; and specifying the encoded component in a bitstream.
[0235] Clause 13F. The method according to Clause 12F, wherein performing psychoacoustic audio coding includes performing psychoacoustic audio coding according to the at least one pair of AptX compression algorithms for the plurality of related components.
[0236] Clause 14F. Any method pursuant to any combination of Clauses 12F and 13F, wherein performing psychoacoustic audio coding comprises performing psychoacoustic audio coding on at least one pair of related components of the plurality of related components to obtain the encoded component.
[0237] Clause 15F. Any combination of methods according to Clauses 12F-14F, wherein performing correlation includes performing correlation individually for the plurality of background components to obtain a plurality of correlated background components, and wherein performing psychoacoustic audio coding includes performing psychoacoustic audio coding for at least one pair of the plurality of background components.
[0238] Clause 16F. Any combination of Clauses 12F-15F, wherein performing correlation includes individually performing correlation on the plurality of foreground audio signals to obtain a plurality of correlated foreground audio signals of the plurality of correlated components, and wherein performing psychoacoustic audio coding includes performing psychoacoustic audio coding on at least one pair of the plurality of correlated foreground audio signals.
[0239] Clause 17F. Any combination of methods according to Clauses 12F-16F, wherein performing correlation includes performing correlation with at least one of the plurality of background components and at least one of the plurality of foreground audio signals to obtain at least one pair of the plurality of correlated components.
[0240] Clause 18F. Any combination of methods pursuant to Clauses 12F-17F further includes: reordering one or more of the plurality of background components and the plurality of foreground audio signals in the bitstream based on the correlation; and specifying in the bitstream an indication of how one or more of the plurality of background components of the plurality of foreground audio signals are reordered in the bitstream.
[0241] Clause 19F. Any combination of Clauses 12F-18F, wherein performing the spatial audio coding comprises performing a linearly reversible transform on the scene-based audio data to obtain the plurality of foreground audio signals and the corresponding plurality of spatial components.
[0242] Clause 20F. Any combination of methods according to Clauses 12F-19F, wherein the scene-based audio data includes higher-order high-fidelity stereo coefficients corresponding to orders greater than zero.
[0243] Clause 21F. Any combination of methods according to Clauses 12F-20F, wherein the scene-based audio data includes audio data defined in the spherical harmonic domain.
[0244] Clause 22F. Any combination of Clauses 12F-21F, wherein each of the plurality of foreground audio signals includes a foreground audio signal defined in the spherical harmonic domain, and wherein each of the corresponding plurality of spatial components includes a spatial component defined in the spherical harmonic domain.
[0245] Clause 23F. An apparatus configured to encode scene-based audio data, the apparatus comprising: means for performing spatial audio coding on the scene-based audio data to obtain a plurality of background components, a plurality of foreground audio signals, and a plurality of corresponding spatial components of a sound field represented by the scene-based audio data, each of the plurality of spatial components defining spatial characteristics of a corresponding foreground audio signal in the plurality of foreground audio signals; means for correlating one or more of the plurality of background components and the plurality of foreground audio signals to obtain a plurality of correlated components; means for performing psychoacoustic audio coding on one or more of the plurality of correlated components to obtain an encoded component; and means for specifying the encoded component in a bitstream.
[0246] Clause 24F. A device according to Clause 23F, wherein the component for performing psychoacoustic audio coding includes performing psychoacoustic audio coding according to at least one pair of AptX compression algorithms for the plurality of related components.
[0247] Clause 25F. Any device pursuant to any combination of Clauses 23F and 24F, wherein the component for performing psychoacoustic audio coding includes a component for performing psychoacoustic audio coding on at least one pair of related components of the plurality of related components to obtain the encoded component.
[0248] Clause 26F. Any combination of Clauses 23F-25F, wherein the component for performing correlation includes components for individually performing correlation against the plurality of background components to obtain a plurality of correlated background components, and wherein the component for performing psychoacoustic audio coding includes components for performing psychoacoustic audio coding for at least one pair of the plurality of background components.
[0249] Clause 27F. Any combination of Clauses 23F-26F, wherein the means for performing correlation includes means for individually performing correlation with respect to the plurality of foreground audio signals to obtain a plurality of correlated foreground audio signals of the plurality of correlated components, and wherein the means for performing psychoacoustic audio coding includes means for performing psychoacoustic audio coding for at least one pair of the plurality of correlated foreground audio signals.
[0250] Clause 28F. Any device pursuant to any combination of Clauses 23F-27F, wherein the component for performing correlation includes a component that performs correlation with at least one of the plurality of background components and at least one of the plurality of foreground audio signals to obtain at least one pair of the plurality of correlated components.
[0251] Clause 29F. Any device pursuant to any combination of Clauses 23F-28F further includes: means for reordering one or more of the plurality of background components and the plurality of foreground audio signals in the bitstream based on the correlation; and means for specifying in the bitstream an indication of how the plurality of background components of the plurality of foreground audio signals are reordered in the bitstream.
[0252] Clause 30F. Any combination of Clauses 23F-29F, wherein the apparatus for performing the spatial audio coding includes components for performing a linearly reversible transformation on the scene-based audio data to obtain the plurality of foreground audio signals and the corresponding plurality of spatial components.
[0253] Clause 31F. Any device according to a combination of Clauses 23F-30F, wherein the scene-based audio data includes higher-order high-fidelity stereo sound coefficients corresponding to orders greater than zero.
[0254] Clause 32F. Any device pursuant to any combination of Clauses 23F-31F, wherein the scene-based audio data includes audio data defined in the spherical harmonic domain.
[0255] Clause 33F. Any device according to any combination of Clauses 23F-32F, wherein each of the plurality of foreground audio signals includes a foreground audio signal defined in the spherical harmonic domain, and wherein each of the corresponding plurality of spatial components includes a spatial component defined in the spherical harmonic domain.
[0256] Clause 34F. A non-transitory computer-readable storage medium having instructions thereon, which, when executed, cause one or more processors to: perform spatial audio coding on scene-based audio data to obtain a plurality of background components, a plurality of foreground audio signals, and a plurality of corresponding spatial components of a sound field represented by the scene-based audio data, each of the plurality of spatial components defining a spatial characteristic of a corresponding foreground audio signal among the plurality of foreground audio signals; correlate the plurality of background components and one or more of the plurality of foreground audio signals to obtain a plurality of correlated components; perform psychoacoustic audio coding on one or more of the plurality of correlated components to obtain an encoded component; and specify the encoded component in a bitstream.
[0257] Clause 1G. An apparatus configured to decode a bitstream representing scene-based audio data, the apparatus comprising: a memory configured to store the bitstream, the bitstream including a plurality of encoded correlated components of a sound field represented by the scene-based audio data; and one or more processors configured to: perform psychoacoustic audio decoding for one or more of the plurality of encoded correlated components to obtain the plurality of correlated components; obtain from the bitstream an indication of how one or more of the plurality of correlated components are reordered in the bitstream; reorder the plurality of correlated components based on the indication to obtain a plurality of reordered components; and reconstruct the scene-based audio data based on the plurality of reordered components.
[0258] Clause 2G. A device according to Clause 1G, wherein the one or more processors are configured to perform psychoacoustic audio decoding on the plurality of encoded related components according to the AptX compression algorithm.
[0259] Clause 3G. A device according to any combination of Clauses 1G and 2G, wherein the one or more processors are configured to perform psychoacoustic audio decoding for at least one pair of the plurality of encoded related components to obtain the plurality of related components.
[0260] Clause 4G. A device according to any combination of Clauses 1G to 3G, wherein the one or more processors are configured to individually reorder the multiple related background components based on the instruction to obtain multiple reordered background components of the multiple reordered components.
[0261] Clause 5G. A device according to any combination of Clauses 1G-4G, wherein the one or more processors are configured to individually reorder the plurality of related foreground audio signals of the plurality of related components based on the instruction to obtain a plurality of reordered foreground audio signals of the plurality of reordered components.
[0262] Clause 6G. Any device according to a combination of Clauses 1G-5G, wherein the plurality of associated components includes a background component associated with the foreground audio signal.
[0263] Clause 7G. Any device according to any combination of Clauses 1G-6G, wherein the scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than 1.
[0264] Clause 8G. Any device according to any combination of Clauses 1G-6G, wherein the scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than zero.
[0265] Clause 9G. Any device according to a combination of Clauses 1G-6G, wherein the scene-based audio data includes audio data defined in the spherical harmonic domain.
[0266] Clause 10G. A device according to any combination of Clauses 1G-9G, wherein the one or more processors are further configured to render the scene-based audio data to one or more speaker feeds, and wherein the device further includes a speaker configured to reproduce the sound field represented by the scene-based audio data based on the speaker feed.
[0267] Clause 11G. A method for decoding a bitstream representing scene-based audio data, the method comprising: obtaining a plurality of encoded related components from the bitstream; performing psychoacoustic audio decoding on one or more of the plurality of encoded related components to obtain the plurality of related components; obtaining from the bitstream an indication of how one or more of the plurality of related components are reordered in the bitstream; reordering the plurality of related components based on the indication to obtain a plurality of reordered components; and reconstructing the scene-based audio data based on the plurality of reordered components.
[0268] Clause 12G. According to the method of Clause 11G, performing psychoacoustic audio decoding includes performing psychoacoustic audio decoding on the plurality of encoded related components according to the AptX compression algorithm.
[0269] Clause 13G. Any method according to a combination of Clauses 11G and 12G, wherein performing psychoacoustic audio decoding includes performing psychoacoustic audio decoding on at least one pair of the plurality of encoded related components to obtain the plurality of related components.
[0270] Clause 14G. Any combination of Clauses 11G-13G, wherein reordering the plurality of related components includes reordering the plurality of related background components individually based on the instruction to obtain a plurality of reordered background components of the plurality of reordered components.
[0271] Clause 15G. Any combination of Clauses 11G-14G, wherein reordering the plurality of related components comprises reordering the plurality of related foreground audio signals individually based on the instruction to obtain a plurality of reordered foreground audio signals of the plurality of reordered components.
[0272] Clause 16G. Any combination of methods according to Clauses 11G-15G, wherein the plurality of related components includes a background component related to the foreground audio signal.
[0273] Clause 17G. Any combination of Clauses 11G-16G, wherein the scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than 1.
[0274] Clause 18G. Any combination of Clauses 11G-16G, wherein the scene-based audio data includes high-order high-fidelity stereo coefficients corresponding to orders greater than zero.
[0275] Clause 19G. Any combination of methods according to Clauses 11G-16G, wherein the scene-based audio data includes audio data defined in the spherical harmonic domain.
[0276] Clause 20G. Any combination of methods according to Clauses 11G-19G further includes: rendering the scene-based audio data to one or more speaker feeds; and outputting the one or more speaker feeds to one or more speakers.
[0277] Clause 21G. A component of an apparatus configured to decode a bitstream representing scene-based audio data, the apparatus comprising: components for obtaining a plurality of encoded correlated components from the bitstream; components for performing psychoacoustic audio decoding on one or more of the plurality of encoded correlated components to obtain the plurality of correlated components; components for obtaining from the bitstream an indication of how to reorder the one or more of the plurality of correlated components in the bitstream; components for reordering the plurality of correlated components based on the indication to obtain a plurality of reordered components; and components for reconstructing the scene-based audio data based on the plurality of reordered components.
[0278] Clause 22G. According to Clause 21G, the apparatus for performing psychoacoustic audio decoding includes apparatus for performing psychoacoustic audio decoding on the plurality of encoded related components according to the AptX compression algorithm.
[0279] Clause 23G. Any device pursuant to any combination of Clauses 21G and 22G, wherein the component for performing psychoacoustic audio decoding includes a component for performing psychoacoustic audio decoding on at least one pair of the plurality of encoded related components to obtain the plurality of related components.
[0280] Clause 24G. Any combination of Clauses 21G-23G, wherein the component for reordering the plurality of related components includes a component for individually reordering the plurality of related background components based on the instruction to obtain a plurality of reordered background components of the plurality of reordered components.
[0281] Clause 25G. Any device according to any combination of Clauses 21G-24G, wherein the component for reordering the plurality of related components includes a component for individually reordering the plurality of related foreground audio signals based on the instruction to obtain a plurality of reordered foreground audio signals of the plurality of reordered components.
[0282] Clause 26G. Any device according to a combination of Clauses 21G-25G, wherein the plurality of related components includes a background component related to the foreground audio signal.
[0283] Clause 27G. Any device according to a combination of Clauses 21G-26G, wherein the scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than 1.
[0284] Clause 28G. Any device according to any combination of Clauses 21G to 26G, wherein the scene-based audio data includes higher-order high-fidelity stereo sound coefficients corresponding to values greater than zero.
[0285] Clause 29G. Any device according to a combination of Clauses 21G-26G, wherein the scene-based audio data includes audio data defined in the spherical harmonic domain.
[0286] Clause 30G. Any device according to any combination of Clauses 21G-29G further includes: means for rendering the scene-based audio data to one or more speaker feeds; and means for feeding the one or more speaker outputs to one or more speakers.
[0287] Clause 31G. A non-transitory computer-readable storage medium having instructions thereon, which, when executed, cause one or more processors to: obtain a plurality of encoded correlated components from a bitstream representing scene-based audio data; perform psychoacoustic audio decoding on one or more of the plurality of encoded correlated components to obtain a plurality of correlated components; obtain from the bitstream an indication of how one or more of the plurality of correlated components are reordered in the bitstream; reorder the plurality of correlated components based on the indication to obtain a plurality of reordered components; and reconstruct the scene-based audio data based on the plurality of reordered components.
[0288] In some contexts, such as broadcast contexts, the audio encoding device may be split into a spatial audio encoder and a psychoacoustic audio encoder 26 (which may also be referred to as "perceptual audio encoder 26"), the spatial audio encoder performing intermediate compression for a high-fidelity stereo representation including gain control, and the psychoacoustic audio encoder performing perceptual audio compression to reduce data redundancy between gain-normalized transmission channels.
[0289] Furthermore, the foregoing techniques can be implemented for any number of different contexts and audio ecosystems, and should not be limited to any of the contexts or audio ecosystems described above. Several exemplary environments are described below, but these techniques should be limited to these exemplary environments. An example audio ecosystem may include audio content, film studios, music studios, game audio studios, channel-based audio content, codec engines, game audio systems, game audio codec / rendering engines, and delivery systems.
[0290] Film studios, music studios, and game audio studios can receive audio content. In some examples, the audio content can represent the output of an acquisition. Film studios can, for example, output channel-based audio content using a digital audio workstation (DAW) (e.g., in 2.0, 5.1, and 7.1). Music studios can, for example, output channel-based audio content using a DAW (e.g., 2.0 and 5.1). In either case, a codec engine can receive and encode channel-based audio content based on one or more codecs (e.g., AAC, AC3, Dolby True HD, Dolby Digital Plus, and DTS master audio) for delivery system output. Game audio studios can, for example, output one or more game audio stems using a DAW. Game audio codec / rendering engines can encode and / or render audio stems into channel-based audio content for delivery system output. Another exemplary scenario for executable technologies includes an audio ecosystem, which may include broadcast-recorded audio objects, professional audio systems, acquisition on consumer devices, high-fidelity stereo reproduction audio formats, on-device rendering, consumer audio, TV and accessories, and automotive audio systems.
[0291] Broadcast recordings, professional audio systems, and consumer devices can all encode and decode their outputs using a high-fidelity stereo reproducible audio format. In this way, audio content can be encoded and decoded into a single representation using a high-fidelity stereo reproducible audio format, and this single representation can be played back using on-device rendering, consumer audio, TV and accessories, and automotive audio systems. In other words, a single representation of audio content can be played back at a general-purpose audio playback system, such as audio playback system 16 (i.e., as opposed to specific configurations requiring such as 5.1, 7.1, etc.).
[0292] Other examples of contexts in which the technology can be implemented include an audio ecosystem, which may include acquisition and playback elements. Acquisition elements may include wired and / or wireless acquisition devices (e.g., Eigen microphones), on-device surround sound acquisition, and mobile devices (e.g., smartphones and tablets). In some examples, wired and / or wireless acquisition devices may be coupled to mobile devices via one or more wired and / or wireless communication channels.
[0293] According to one or more techniques disclosed herein, a mobile device can be used to acquire a sound field. For example, the mobile device can acquire a sound field via wired and / or wireless acquisition devices and / or on-device surround sound acquisition (e.g., multiple microphones integrated into the mobile device). The mobile device can then encode the acquired sound field into high-fidelity stereo reproduction coefficients for playback by one or more playback elements. For example, a user of the mobile device can record a live event (e.g., a meeting, concert, etc.) (acquiring the sound field of the live event) and encode the recording into high-fidelity stereo reproduction coefficients.
[0294] Mobile devices can also utilize one or more playback elements to reproduce the sound field of a high-fidelity stereo codec. For example, a mobile device can decode the sound field of a high-fidelity stereo codec and output a signal to one or more playback elements, which causes the one or more playback elements to reconstruct the sound field. As an example, a mobile device can utilize wireless and / or wireless communication channels to output a signal to one or more speakers (e.g., speaker arrays, soundbars, etc.). As another example, a mobile device can utilize docking solutions to output a signal to one or more docking stations and / or one or more docked speakers (e.g., sound systems in smart cars and / or homes). As yet another example, a mobile device can utilize headphone rendering to output a signal to a set of headphones, for example, to produce realistic binaural sound.
[0295] In some examples, a specific mobile device may acquire a 3D sound field and play back the same 3D sound field at a later time. In some examples, a mobile device may acquire a 3D sound field, encode that 3D sound field into a high-fidelity stereo reproduction, and transmit the encoded 3D sound field to one or more other devices (e.g., other mobile devices and / or other non-mobile devices) for playback.
[0296] Another context in which this technology can be implemented includes the audio ecosystem, which may include audio content, game studios, encoded audio content, rendering engines, and delivery systems. In some examples, a game studio may include one or more DAWs that support editing of high-fidelity stereo signals. For example, one or more DAWs may include high-fidelity stereo reproduction plugins and / or tools that can be configured to operate (e.g., work with) one or more game audio systems. In some examples, a game studio may output a new stem format that supports high-fidelity stereo reproduction. In any case, a game studio may output encoded audio content to a rendering engine that renders the sound field for playback by a delivery system.
[0297] These techniques can also be performed on exemplary audio acquisition devices. For example, these techniques can be performed on an Eigen microphone, which may include multiple microphones collectively configured to record a 3D sound field. In some examples, the multiple microphones of the Eigen microphone may be located on the surface of a substantially spherical sphere with a radius of approximately 4 cm. In some examples, audio encoding device 20 may be integrated into the Eigen microphone to output bitstream 21 directly from the microphone.
[0298] Another exemplary audio acquisition context may include a production vehicle, which can be configured to receive signals from one or more microphones, such as one or more Eigen microphones. The production vehicle may also include an audio encoder, such as... Figure 1 Spatial audio encoding device 24.
[0299] In some cases, the mobile device may also include multiple microphones collectively configured to record a 3D sound field. In other words, these multiple microphones may have X, Y, Z diversity. In some examples, the mobile device may include microphones that can be rotated to provide X, Y, Z diversity for one or more other microphones on the mobile device. The mobile device may also include an audio encoder, such as... Figure 1 The audio encoder 22.
[0300] The ruggedized video capture device can also be configured to record 3D sound fields. In some examples, the ruggedized video capture device can be attached to the helmet of a user participating in the activity. For example, the ruggedized video capture device can be attached to a user's helmet during whitewater rafting. In this way, the ruggedized video capture device can capture 3D sound fields representing actions around the user (e.g., water hitting behind the user, another raftsman speaking in front of the user, etc.).
[0301] These techniques can also be performed on accessory-enhanced mobile devices that can be configured to record 3D sound fields. In some examples, the mobile device can be similar to the mobile devices discussed above, with one or more accessories added. For example, an Eigen microphone can be attached to the aforementioned mobile device to form an accessory-enhanced mobile device. In this way, the accessory-enhanced mobile device can capture a higher quality version of the 3D sound field than using only the sound acquisition components integrated into the accessory-enhanced mobile device.
[0302] Exemplary audio playback devices that can perform various aspects of the techniques described in this disclosure are also discussed below. According to one or more techniques of this disclosure, speakers and / or soundbars can be arranged in any arbitrary configuration while still reproducing a 3D sound field. Furthermore, in some examples, headphone playback devices can be coupled to decoder 32 (this refers to...) via a wired or wireless connection. Figure 1 (Another approach to the audio decoding device 32). According to one or more techniques of this disclosure, a single universal representation of the sound field can be used to render the sound field on any combination of loudspeakers, sound bars, and headphone playback devices.
[0303] Several different exemplary audio playback environments may also be suitable for performing various aspects of the techniques described in this disclosure. For example, a 5.1 speaker playback environment, a 2.0 (e.g., stereo) speaker playback environment, a 9.1 speaker playback environment with full-height front speakers, a 22.2 speaker playback environment, a 16.0 speaker playback environment, a car speaker playback environment, and a mobile device with an earphone playback environment may be suitable environments for performing various aspects of the techniques described in this disclosure.
[0304] According to one or more techniques of this disclosure, a sound field can be rendered in any of the aforementioned playback environments using a single universal representation of the sound field. Additionally, the techniques of this disclosure enable the renderer to render a sound field from a universal representation for playback in playback environments different from those described above. For example, if the design precludes proper speaker placement for a 7.1 speaker playback environment (e.g., if placing a right surround speaker is not possible), then the techniques of this disclosure enable the renderer to compensate with six additional speakers to achieve playback in a 6.1 speaker playback environment.
[0305] Furthermore, users can watch sports games while wearing headphones. According to one or more techniques disclosed herein, a 3D sound field of a sports game can be acquired (e.g., one or more Eigen microphones can be placed in and / or around a baseball field), a high-fidelity stereo sound reproduction factor corresponding to the 3D sound field can be obtained and transmitted to a decoder, which can reconstruct the 3D sound field based on the high-fidelity stereo sound reproduction factor and output the reconstructed 3D sound field to a renderer, which can obtain an indication of the type of playback environment (e.g., headphones) and render the reconstructed 3D sound field into a signal that causes the headphones to output a presentation of the 3D sound field of the sports game.
[0306] In each of the various examples described above, it should be understood that the audio encoding device 22 may perform a method or otherwise include components for performing each step of the method to which the audio encoding device 22 is configured. In some cases, the device may include one or more processors. In some cases, the one or more processors may represent a dedicated processor configured by instructions stored in a non-transitory computer-readable storage medium. In other words, various aspects of the technology in each of the encoding example sets can provide a non-transitory computer-readable storage medium on which the instructions are stored above, which, when executed, cause one or more processors to perform the method to which the audio encoding device 20 is configured.
[0307] In one or more examples, the described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions can be stored or sent as one or more instructions or code to a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium. A data storage medium can be any available medium accessible by one or more computers or processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products can include computer-readable media.
[0308] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that can be accessed by a computer. However, it should be understood that computer-readable storage media and data storage media do not include links, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used in this application, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks generally reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0309] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be implemented entirely within one or more circuit or logic elements.
[0310] The techniques disclosed herein can be implemented in a variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to highlight functional aspects of a device configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Rather, as described above, various units can be combined in a codec hardware unit, or provided by a set of interoperable hardware units including one or more processors as described above, combined with suitable software and / or firmware.
[0311] In addition, as used in this article, “A and / or B” means “A or B”, or “A and B”.
[0312] Various aspects of the technology have been described. These and other aspects of the technology are within the scope of the following claims.
Claims
1. A device configured to encode scene-based audio data, the device comprising: A memory is configured to store the scene-based audio data; as well as One or more processors are configured as follows: Spatial audio coding is performed on the scene-based audio data to obtain multiple background components, multiple foreground audio signals, and multiple corresponding spatial components of the sound field represented by the scene-based audio data, each of the multiple spatial components defining the spatial characteristics of the corresponding foreground audio signal in the multiple foreground audio signals; Perform correlation between at least one of the plurality of background components and at least one of the plurality of foreground audio signals to obtain a plurality of correlated components; Psychoacoustic audio coding is performed on one or more of the plurality of related components to obtain the encoded components; as well as Specify the encoded component in the bitstream.
2. The device according to claim 1, wherein, The one or more processors are configured to perform psychoacoustic audio coding according to a compression algorithm for at least one pair of related components among the plurality of related components.
3. The device according to claim 1, wherein, The one or more processors are configured to perform psychoacoustic audio coding on at least one pair of related components among the plurality of related components to obtain encoded components.
4. The device according to claim 1, wherein, The one or more processors are further configured to: Perform correlation analysis on each of the multiple background components individually to obtain multiple correlated background components; as well as Psychoacoustic audio coding is performed for at least one pair of the plurality of background components.
5. The device according to claim 1, wherein, The one or more processors are further configured to: Correlation is performed individually on the multiple foreground audio signals to obtain multiple correlated foreground audio signals of the multiple correlated components; as well as Psychoacoustic audio coding is performed on at least one pair of the plurality of associated foreground audio signals.
6. The device according to claim 1, wherein, The one or more processors are further configured to: Based on the correlation, one or more of the plurality of background components and the plurality of foreground audio signals in the bitstream are reordered; as well as The bitstream specifies an indication of how one or more of the plurality of background components and the plurality of foreground audio signals are reordered in the bitstream.
7. The device according to claim 1, wherein, The one or more processors are configured to perform a linear reversible transformation on the scene-based audio data to obtain the plurality of foreground audio signals and the corresponding plurality of spatial components.
8. The device according to claim 1, wherein, The scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than zero.
9. The device according to claim 1, wherein, The scene-based audio data includes audio data defined in the spherical harmonic domain.
10. The device according to claim 1, in, Each of the plurality of foreground audio signals includes a foreground audio signal defined in the spherical harmonic domain, and Each of the corresponding plurality of spatial components includes a spatial component defined in the spherical harmonic domain.
11. The device according to claim 1, in, At least one of the plurality of background components includes two or more background components, and Wherein, at least one of the plurality of foreground audio signals includes two or more foreground audio signals.
12. A method for encoding scene-based audio data, the method comprising: Spatial audio coding is performed on the scene-based audio data to obtain multiple background components, multiple foreground audio signals, and multiple corresponding spatial components of the sound field represented by the scene-based audio data, each of the multiple spatial components defining the spatial characteristics of the corresponding foreground audio signal in the multiple foreground audio signals; Perform correlation between at least one of the plurality of background components and at least one of the plurality of foreground audio signals to obtain a plurality of correlated components; Psychoacoustic audio coding is performed on one or more of the plurality of related components to obtain the encoded components; as well as Specify the encoded component in the bitstream.
13. An apparatus configured to decode a bitstream representing scene-based audio data, the apparatus comprising: A memory is configured to store the bitstream, the bitstream comprising multiple encoded related components of a sound field represented by the scene-based audio data; as well as One or more processors are configured as follows: Perform psychoacoustic audio decoding on one or more of the plurality of encoded related components to obtain the plurality of related components; Obtain from the bitstream an indication of how one or more of the plurality of related components are reordered in the bitstream, wherein the plurality of related components include a background component related to the foreground audio signal; Based on the instruction, the plurality of related components are reordered to obtain a plurality of reordered components; and The scene-based audio data is reconstructed based on the multiple reordered components.
14. The device according to claim 13, wherein, The one or more processors are configured to perform psychoacoustic audio decoding on the plurality of encoded related components according to a decompression algorithm.
15. The device according to claim 13, wherein, The one or more processors are configured to perform psychoacoustic audio decoding on at least one pair of encoded correlated components to obtain the plurality of correlated components.
16. The device according to claim 13, wherein, The one or more processors are configured to individually reorder multiple related background components of the plurality of related components based on the indication to obtain multiple reordered background components of the plurality of reordered components.
17. The device according to claim 13, wherein, The one or more processors are configured to individually reorder multiple related foreground audio signals of the multiple related components based on the indication to obtain multiple reordered foreground audio signals of the multiple reordered components.
18. The device according to claim 13, wherein, The scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than one.
19. The device according to claim 13, wherein, The scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than zero.
20. The device according to claim 13, wherein, The scene-based audio data includes audio data defined in the spherical harmonic domain.
21. The device according to claim 13, in, The one or more processors are further configured to render the scene-based audio data to one or more speaker feeds, and The device also includes a speaker configured to reproduce a sound field represented by the scene-based audio data based on the speaker feed.
22. The device according to claim 13, in, The background component includes one or more background components, and The foreground audio signal includes one or more foreground audio signals.
23. A method for decoding a bitstream representing scene-based audio data, the method comprising: Multiple encoded related components are obtained from the bitstream; Perform psychoacoustic audio decoding on one or more of the plurality of encoded related components to obtain the plurality of related components; Obtain from the bitstream an indication of how one or more of the plurality of related components are reordered in the bitstream, wherein the plurality of related components include a background component related to the foreground audio signal; Based on the instruction, the plurality of related components are reordered to obtain a plurality of reordered components; and The scene-based audio data is reconstructed based on the multiple reordered components.
24. The method according to claim 23, wherein, Performing psychoacoustic audio decoding includes performing psychoacoustic audio decoding on the plurality of encoded related components according to a decompression algorithm.
25. The method according to claim 23, wherein, Performing psychoacoustic audio decoding includes performing psychoacoustic audio decoding on at least one pair of encoded related components from the plurality of encoded related components to obtain the plurality of related components.
26. The method according to claim 23, wherein, Reordering the plurality of related components includes reordering the plurality of related background components individually based on the instruction to obtain the plurality of reordered background components of the plurality of reordered components.
27. The method according to claim 23, wherein, Reordering the plurality of related components includes reordering the plurality of related foreground audio signals of the plurality of related components individually based on the indication to obtain the plurality of reordered foreground audio signals of the plurality of reordered components.
28. The method according to claim 23, wherein, The scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than one.
29. The method according to claim 23, wherein, The scene-based audio data includes high-order high-fidelity stereo sound coefficients corresponding to orders greater than zero.
30. A computer-readable medium having program code recorded thereon, wherein the program code is executable by one or more processors of a device to cause the one or more processors to perform the method according to claim 12.
31. A computer-readable medium having program code recorded thereon, wherein the program code is executable by one or more processors of a device to cause the one or more processors to perform the method according to any one of claims 23-29.
Citation Information
Patent Citations
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US10405126B2
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US20190007781A1
Compensating for error in decomposed representations of sound fields
CN105264598A
Reducing correlation between higher order ambisonic (HOA) background channels
CN106663433A