stereo-based immersive coding (stic)

By using encoding and decoding technology based on dual-channel stereo signals and directional parameters, the problem of audio quality degradation under bandwidth limitations is solved, achieving a high-quality immersive audio experience at low bit rates and reducing spectral distortion.

CN115989682BActive Publication Date: 2026-01-02APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180052259.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-27
Filing Date
2021-08-20
Publication Date
2026-01-02
Estimated Expiration
2041-08-20

AI Technical Summary

Technical Problem

Existing technologies suffer from limited bandwidth when transmitting immersive audio content, leading to a decline in audio quality and failing to effectively utilize auditory perception characteristics to reduce data size.

Method used

By using encoding and decoding techniques based on two-channel stereo signals and directional parameters, an immersive audio experience is recreated. Spatial rendering is achieved using time-frequency patching and weighting factors, reducing the correlation between channel pairs, lowering the bit rate, and reducing spectral distortion.

Benefits of technology

While reducing the bitrate, it improves the audio quality of immersive audio content, reduces distortion caused by spatial rendering, and provides a stable audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115989682B_ABST
    Figure CN115989682B_ABST
Patent Text Reader

Abstract

An audio codec represents an immersive signal through binaural stereo signals and directional parameters, the binaural stereo signals being a stereo rendering of the immersive signal. The directional parameters can recreate the perceived location of a dominant sound based on a perceptual model describing the direction of a virtual loudspeaker pair. Audio processing at the decoder can be performed on the stereo signals in the frequency domain of multiple channel pairs using time-frequency tiles. Spatial positioning of the audio signal can use a panning method, in particular by applying weights to the time-frequency tiles of the stereo signals for each output channel pair. The weights for the time-frequency tiles can be derived based on the directional parameters, an analysis of the stereo signals, and an output channel layout. The weights can be used to adaptively process the time-frequency tiles using decorrelators to reduce or minimize spectral distortion due to spatial rendering.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 071,149, filed August 27, 2020, the disclosure of which is incorporated by reference herein in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates to the field of audio communication; more specifically, to a digital signal processing method designed to convey immersive audio content using stereo signals. Other aspects are also described. BACKGROUND

[0004] Consumer electronics devices are providing increasingly complex and ever-improving digital audio encoding and decoding capabilities. Traditionally, audio content has been produced, distributed, and consumed primarily using a two-channel stereo format that provides left and right audio channels. Recent market developments aim to provide a more immersive listener experience using richer audio formats (e.g., Dolby Atmos or MPEG-H) that support multi-channel audio, object-based audio, and / or Ambisonics.

[0005] The conveyance of immersive audio content is associated with greater bandwidth needs, i.e., greater data rates are needed for streaming and downloading compared to for stereo content. If bandwidth is limited, techniques are needed that can reduce the size of the audio data while maintaining the best possible audio quality. A common approach to reduce bandwidth in perceptual audio coding is to exploit perceptual properties of the human ear to maintain audio quality. For example, at the lowest bitrates, audio coding can exploit parametric methods to bit-rate-efficiently encode certain sound features so that they can be approximately recreated in the decoder. An example of parametric surround audio coding is MPEG Surround or Binaural Cue Coding (BCC), which can use spatial parameters to recreate a multi-channel audio signal from a mono audio signal. Other audio coding and decoding (codec) techniques are also needed to convey richer and more immersive audio content using limited bandwidth. SUMMARY

[0006] A new immersive audio codec is disclosed that recreates an immersive audio experience based on a two-channel stereo signal and direction parameters. The stereo signal is a high-quality stereo rendering of the immersive audio signal, and the direction parameters can be based on a perceptual model that derives parameters describing the direction of a perceived dominant sound. The immersive audio signal can include multi-channel audio, audio objects, or higher-order ambisonic (HOA) that describes a soundfield based on spherical harmonics. For example, when the immersive audio signal is a multi-channel input of more than two channels, it can be down-mixed to a stereo signal. When the immersive audio signal represents audio objects or HOA components, the objects or HOA components can be rendered to a stereo signal. The stereo signal and the direction parameters can be encoded by an encoder and transmitted to a decoder for reconstruction and playback.

[0007] At the decoder, the decoded stereo signal can be converted from time domain to frequency domain and separated into time-frequency tiles. The left and right signals of the time-frequency tiles can be processed in parallel by multiple processing units, each associated with a pair of playback channels or speakers. Weighting factors can be applied to the tiles to generate corresponding weighted time-frequency tiles for the output channel pair. The weighting factors can be controlled to create a perceived direction from which the audio signals of the time-frequency tiles will be heard in the multi-channel playback system, given a playback channel layout. The direction parameters received from the encoder can represent the direction of a perceived dominant sound in subbands of the time-frequency tiles, and the direction parameters can be used by the decoder to control the weighting factors.

[0008] In one aspect, the decoder can control the weighting factors based on an analysis of the stereo signal and the direction parameters to reduce correlation between channel pairs. De-correlation can be applied to reduce comb-filtering effects that can cause large image shifts of the perceived audio signal when a listener moves. The effects can be pronounced in audio signals with smooth envelopes and high prediction gain. The decoder can analyze the stereo signal and the direction parameters to generate weighting factors for de-correlation and estimate an amount of de-correlation for each time-frequency tile. In one aspect, to mitigate distortions due to spatial rendering, such as unstable images caused by concurrent sources present in different directions or time smearing of attack of transient signals, the decoder can estimate temporal fluctuations of the perceived direction of dominance in subbands of the time-frequency tiles to control the generation of the weighting factors.

[0009] After applying the weighting factors to the spatially rendered time- frequency tiles of the channel pairs, the weighted time-frequency tiles are merged to convert the left and right signals of each channel pair from the frequency domain back to the time domain. The time-domain signals for these channel pairs can be combined to generate signals for the speakers of a multi-channel playback system. In one aspect, the stereo signal can be used as a fallback audio signal for systems that cannot decode the direction parameters, have only a stereo playback system, or whose stereo signal is preferred for headphone playback.

[0010] Advantageously, to reduce the bit rate, aspects of the disclosure reduce the number of audio channels to be transmitted to two channels. For the direction parameters, the stereo signal uses only a small amount of side information, far below the bit rate required for a single audio channel. Signal processing is performed based on these direction parameters and an analysis of the stereo signal to reduce or minimize spectral distortion due to spatial rendering using techniques such as temporal smoothing and decorrelation of the weighting factors. The audio quality of immersive audio content can be enhanced while achieving a reduction in bit rate.

[0011] In one aspect, a method for encoding audio content is disclosed. The method includes generating a two-channel stereo signal from audio content, such as an immersive audio signal. The method also includes generating direction parameters based on the audio content. The direction parameters describe optimal directions of pairs of virtual loudspeakers to recreate a perceived dominant sound location of the audio content in a plurality of frequency subbands. The method further includes transmitting the two-channel stereo signal and the direction parameters to a decoding device over a communication channel.

[0012] In one aspect, a method for decoding audio content is disclosed. The method includes receiving a two-channel stereo signal and direction parameters from an encoding device. The direction parameters describe optimal directions of pairs of virtual loudspeakers to recreate a perceived dominant sound location of audio content represented by the two-channel stereo signal in a plurality of frequency subbands. The method also includes generating a plurality of time-frequency tiles of a plurality of channel pairs of a playback system from the two-channel stereo signal. The plurality of time-frequency tiles represent a frequency domain representation of each channel of the two-channel stereo signal in the plurality of frequency subbands. The method further includes generating weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs based on the direction parameters. The method also includes applying the weighting factors to the plurality of time-frequency tiles to spatially render the time-frequency tiles by the plurality of channel pairs of the playback system.

[0013] The above summary does not include an exhaustive list of all aspects of the present application. It is contemplated that all systems and methods encompassed by this application can be practiced with the individual aspects and combinations of the aspects disclosed in the summary above, as well as in the detailed description below and throughout the claims, which follow this patent application. Such combinations have specific advantages that are not specifically recited in the above summary. BRIEF DESCRIPTION OF DRAWINGS

[0014] Aspects of the disclosure are illustrated by way of example and not by way of limitation in the figures of the accompanying drawings in which like references indicate similar elements. It should be noted that "a" or "one" aspect as referred to herein does not necessarily pertain to the same aspect, and that "a" or "one" aspect can refer to at least one. Additionally, to be clear, features of the present disclosure can be shown in a given figure as being part of two separate aspects even though the aspects are shown separately, and that an aspect can include features of the other aspect, and vice versa. Stated differently, for conciseness and clarity, a given figure can show the features of more than one aspect, and vice versa, and these apparent separate aspects can also include features of the other apparent separate aspects.

[0015] Figure 1 is a functional block diagram of a stereo-based immersive audio encoding system according to one aspect of the disclosure.

[0016] Figure 2 depicts top views of five loudspeaker layouts according to one aspect of the disclosure.

[0017] Figure 3 depicts phantom image locations of audio sources perceived from five loudspeaker layouts according to one aspect of the disclosure.

[0018] Figure 4 is a functional block diagram of a stereo-based immersive audio encoding system according to one aspect of the disclosure, including a processing module for reducing or minimizing distortion due to spatial rendering.

[0019] Figure 5 is a functional block diagram of a perceptual model of a stereo-based immersive audio encoding system according to one aspect of the disclosure, the perceptual model for estimating direction parameters.

[0020] Figure 6 is a functional block diagram of a perceptual model of a stereo-based immersive audio encoding system according to one aspect of the disclosure, the perceptual model for estimating direction parameters from channel-based input.

[0021] Figure 7 depicts use of virtual channel pairs for object rendering when a perceptual model of a stereo-based immersive audio encoding system uses azimuth / elevation of virtual channel pairs as metadata according to one aspect of the disclosure.

[0022] Figure 8is a functional block diagram of a decoder process of a channel pair of a stereo-based immersive audio coding system according to an aspect of the disclosure.

[0023] Figure 9 is a functional block diagram of an audio analysis module of a stereo-based immersive audio coding system for adjusting weighting factors according to an aspect of the disclosure.

[0024] Figure 10 is a functional block diagram of a weighting control module for generating weighting factors for time-frequency tiles according to an aspect of the disclosure.

[0025] Figure 11 depicts downmixing of audio channels for multiple sectors of a seven- speaker layout according to an aspect of the disclosure.

[0026] Figure 12 is a functional block diagram of a stereo-based immersive audio coding system for encoding and decoding multiple sections or sectors of a speaker layout according to an aspect of the disclosure.

[0027] Figure 13 is a functional block diagram of a hybrid stereo-based immersive audio coding system according to an aspect of the disclosure that encodes and decodes a single channel such as a center channel independently of other channels encoded and decoded using a STIC system.

[0028] Figure 14 is a flowchart of an encoder-side processing method of a stereo-based immersive audio coding system according to an aspect of the disclosure to generate stereo signals and direction parameters from immersive audio signals.

[0029] Figure 15 is a flowchart of a decoder-side processing method of a stereo-based immersive audio coding system according to an aspect of the disclosure to reconstruct immersive audio signals for a multi-channel playback system. DETAILED DESCRIPTION

[0030] We desire to provide immersive audio content from an audio source to a playback system over a bandwidth-limited transmission channel while maintaining the best possible audio quality. Immersive audio content can include multi-channel audio, audio objects, or spatial audio reconstruction (known as Ambisonics), which describes a soundfield based on spherical harmonics that can be used to recreate a soundfield for playback. Ambisonics can include first-order or higher-order spherical harmonics, also known as higher-order Ambisonics (HOA). Immersive audio content can be rendered to a lower bit rate audio content and spatial parameters can be generated to exploit perceptual properties of the human ear. An encoder can transmit the lower bit rate audio content and spatial parameters over the limited bandwidth channel to allow a decoder to recreate an immersive audio experience.

[0031] Systems and methods are disclosed for immersive audio encoding techniques that recreate an immersive audio experience based on a two-channel stereo signal and directional parameters. Audio processing at a decoder can be performed on left and right signals of a stereo signal using time-frequency tiles in the frequency domain for multiple channel pairs. Directional parameters can indicate an optimal direction of a virtual speaker pair to recreate a perceived dominant sound location for the time-frequency tiles. Spatial positioning of the decoded audio signal can use a panning method in a midplane between channel pairs of a multi-channel playback system, specifically by applying weighting factors to the time-frequency tiles of the stereo signal for each output channel pair. The decoder can derive the weighting factors for the time-frequency tiles based on the directional parameters describing the virtual speaker pair direction, an analysis of the decoded stereo signal, and an output channel layout. These weighting factors can be used to adaptively process the time-frequency tiles using a decorrelator to reduce or minimize spectral distortion due to spatial rendering of the encoding techniques.

[0032] The following description sets forth numerous specific details. It should be understood, however, that aspects of the disclosure might be practiced without these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail in order not to obscure the understanding of this description.

[0033] The terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting of the application. Spatially relative terms, such as "under", "below", "lower", "on", "above", "upper", and the like, can be used herein for ease of description to describe one element or feature's relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device including an element is turned over, elements described as "below" or "beneath" other elements or features would then be oriented "above" the other elements or features. Thus, the exemplary term "below" can encompass both an orientation of above and below. The device can be otherwise oriented (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly.

[0034] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and "comprising", when used in this specification, specify the presence of stated features, steps, operations, elements, or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, or groups thereof.

[0035] The terms "or" and "and / or" as used herein are to be interpreted as inclusive or meaning any one or any combination. Therefore, "A, B or C" or "A, B and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B and C." An exception to this definition will occur only when two elements are directly exclusive with each other according to their definitions. That is, the definition "A or B" does not allow the inclusion of both A and B, unless specifically stated otherwise.

[0036] Figure 1is a functional block diagram of a stereo-based immersive coding (STIC) system according to one aspect of the disclosure. The audio input to the STIC system can include various immersive audio input formats, such as multi-channel audio, audio objects, HOA. It should be appreciated that HOA can also include first order high fidelity stereo ambisonic (FOA). To reduce the data bitrate, a downmixer / renderer module 105 can reduce the audio input to a two-channel stereo signal. In the case of multi-channel input, there can be M channels of a known input channel layout, such as a 7.1.4 layout (7 speakers at mid-plane, 4 speakers at upper plane, 1 low frequency effect (LFE) speaker). The downmixer / renderer module 105 can downmix the multi-channel input, excluding the LFE channel, to a stereo signal. In the case of audio objects, all M objects can first be rendered by the downmixer / renderer 105 to a stereo signal. In the case of HOA, there can be M HOA components, where M depends on the HOA order. The downmixer / renderer 105 can render the HOA signal to a stereo signal. The two-channel stereo signal can be referred to as a right channel signal and a left channel signal.

[0037] The stereo audio signal can be encoded by an encoder of an audio codec 109 to reduce the audio bitrate. The audio codec 109 can use any known encoding and decoding techniques, without further elaboration. A parameter generation module 107 can generate spatial image parameter descriptions of the audio input. The decoder side or receiver of the STIC system uses these spatial image parameters to reconstruct the immersive audio content from the stereo signal. In one aspect, these spatial image parameters can be parameters that describe the best direction of a virtual speaker pair to recreate the location of a perceived dominant sound. In one aspect, the spatial image parameters can be encoded prior to transmission. The encoder side or transmitter of the STIC system can transmit the encoded stereo signal and the spatial image parameters over a bandwidth-limited channel to the decoder side. In one aspect, the bandwidth-limited channel can be a wired or wireless communication medium. In another aspect, the encoder side can encode the stereo signal and the spatial image parameters to reduce or minimize the file size for storage. The decoder side can later retrieve the stored file, which contains the encoded stereo signal and the encoded spatial image parameters for decoding and playback.

[0038] At the decoder side, a decoder of the audio codec 109 can decode the encoded stereo signal. A time-frequency tile separator 111 can convert the decoded stereo signal from the time domain to the frequency domain (such as by a short-time Fourier transform (STFT)) to generate B tiles across the frequency domain. Each tile of the B tiles can represent a frequency subband of the decoded stereo signal for a particular time period. The number of subbands B can be determined by the required spectral resolution. In one aspect, each subband can comprise a grouping of multiple frequency bins from the STFT. In one aspect, the decoded audio signal can be divided into blocks of fixed time periods (also referred to as frame size), represented by B tiles in the frequency domain. The frequency domain representation of the stereo signal can be separated or duplicated into P parallel processing paths, where each processing path can be associated with a pair of playback channels or speakers. Thus, the stereo signal can be separated into PxB time-frequency tiles, each tile representing a subband of the frequency domain representation of the left and right channels of the stereo signal for a pair of playback channels or speakers for a frame duration.

[0039] The time-frequency tile weighting control module 115 can generate weighting factors w(p, b) that are applied to the corresponding PxB tiles of the stereo signal to generate weighted time-frequency tiles for the P output channel pairs. The weighting factors w(p, b) control the spatial rendering to create a perceived direction from which the audio signals of the time-frequency tiles will be heard in the multi-channel playback system given the playback channel layout. The direction parameters received from the encoder can represent the optimal direction of the virtual speaker pair to recreate the perceived location of the dominant sound in the subband of the time-frequency tile, and these direction parameters can be used by the time-frequency tile weighting control module 115 to control the weighting factors w(p, b).

[0040] The time-frequency tile merger module 113 can merge the weighted PxB time-frequency tiles to convert the left and right signals of each output channel pair from the frequency domain back to the time domain. In one aspect, this operation can be the inverse of the operation of the time-frequency tile separator 111. The time-frequency tile merger module 113 can combine the time domain signals for the P output channel pairs to generate the audio signals for the N speakers of the multi-channel playback system. In one aspect, the number of speakers N can not be 2xP.

[0041] Figure 2 A top view of a five-speaker (N = 5) layout of a playback system is depicted in accordance with one aspect of the disclosure. Figure 2 A 5.0 speaker layout is shown in which the five speakers in the median plane are laid out in a circular arrangement in the horizontal plane relative to a listener located in the center. The channel pair as used here refers to the channels assigned to two speakers positioned symmetrically left and right relative to a forward-facing listener. For example, in a 5.0 speaker layout, the left and right channels of the audio signal are assigned to the left and right speakers, respectively, and the center channel is assigned to the center speaker. Figure 2In the middle, the channels assigned to the loudspeaker with p = 3 belong to channel pair 3. For simplicity of the description, a single loudspeaker located in the midplane can have two associated channels added to provide the loudspeaker signal. Thereby, such a loudspeaker is also associated with a channel pair (see, e.g., Fig. 1). Figure 2 In the middle, the channels assigned to the loudspeaker with p = 3 belong to channel pair 3. For simplicity of the description, a single loudspeaker located in the midplane can have two associated channels added to provide the loudspeaker signal. Thereby, such a loudspeaker is also associated with a channel pair (see, e.g., Fig. 1).

[0042] If the weighting factor w(p, b) is set to 0 for all channel pairs except p = 3 (e.g., w(p, b) is set to 1 for p = 3), the audio signal of the time-frequency tile will be directed entirely to channel pair 3, as indicated by the arrow in Figure 1 , and the listener will localize the sound from this direction. By assigning non-zero weighting factors to more than one channel pair, the perceived sound location can be further manipulated. For example, if the weighting factors for channel pairs 2 and 3 have the same value, the sound will be perceived somewhere between the loudspeakers associated with these channel pairs. That is, the source localization in a stereo audio signal is largely based on the so-called phantom image phenomenon. Figure 2

[0043] Figure 3 depicts the phantom image locations of an audio source perceived from the same five loudspeaker (N = 5) layout according to one aspect of the present disclosure. The loudspeaker associated with p = 1 is not shown to avoid obscuring some details depicted in the figure. In Figure 3 , if the two loudspeakers of channel pair 2 (p = 2) emit the same sound, the listener will perceive a phantom image between the two loudspeakers in front. Similarly, if now the same sound signal is emitted by channel pair 3 (p = 3), the listener will perceive a phantom image between the two loudspeakers of channel pair 3. By manipulating the weighting factors for channel pairs 2 and 3, the phantom image location can be shifted to any location between these pairs of loudspeakers.

[0044] ​The same weighting factor can be applied to the left and right signals of a channel pair. Subsequently, the phantom image will remain at the same perceived lateral position as in the stereo downmix signal. Since the dialogue in a movie soundtrack or the lead vocal in a music recording is usually panned to the center position, it can be very important to maintain the perceived position of such a main sound scene element. The spatial positioning of the phantom image of the STIC system includes a panning approach for the decoded stereo signal in the midplane between the channel pairs of a multi-channel playback system. The panning can vary over time and frequency with the support of a tile-based processing based on weighting factors w(p, b) and spatial image parameters. For example, the weighting factors w(p, b) can be derived based on an analysis of the decoded stereo signal and direction parameters describing the direction of a virtual loudspeaker pair to recreate a dominant sound in a subband of the decoded stereo signal. In one aspect, the weighting factors w(p, b) can be used to adaptively process the time-frequency tiles to reduce or minimize spectral distortion due to spatial positioning.

[0045] The synthesis of immersive audio content from a stereo signal using time- frequency tiles as described can achieve the desired spatial positioning, but can also introduce various distortions to the audio playback signal. For example, an unstable image can be perceived when concurrent sources in different directions are present. Distortion can also occur due to time smearing of attacks or transients in the stereo signal. A comb filtering effect can exist when highly correlated signals are generated for multiple output channels. Such effects can cause large image shifts when the listener moves around. Other distortions can include a coloring effect or loudness modulation when the relative magnitude of various frequency components of a wideband sound changes.

[0046] Figure 4 is a functional block diagram of a stereo-based immersive audio encoding system according to one aspect of the disclosure, which includes additional processing modules for reducing or minimizing distortion due to spatial positioning in order to enhance audio quality. The downmixer / renderer 105 and the audio codec 109 can be the same as in Figure 1 and the description of these modules will not be repeated for the sake of brevity.

[0047] The perception model 117 derives parameters that describe the optimal direction of a virtual speaker pair to recreate the perceived location of the dominant sound of the audio input signal. In one aspect, the direction of the virtual speaker pair can be estimated for frequency subbands using a time-frequency tiling. The spectral resolution of the frequency subbands used internally by the perception model 117 for direction estimation can be different (e.g., higher) than the spectral resolution of the frequency subbands used by the time-frequency tiling separator 111 for the decoded stereo signal. The perception model 117 can map the direction of the virtual speaker pair estimated for the internal frequency subbands to the B subbands of the decoded stereo signal. The direction of the virtual speaker pair for each of the B subbands can be given as an azimuth and an elevation angle (in degrees) relative to a default listener position. The azimuth and elevation angles can represent the optimal location of the virtual speaker pair for recreating the dominant sound at the original location. A parametric codec 119 can encode the direction parameters to reduce the data rate for transmission. At the decoder side, a decoder of the parametric codec 119 can decode the received parameters to send the direction parameters to the weighting control module 123. In one aspect, the decoded stereo signal can be used as a fallback audio signal for systems that cannot decode the direction parameters, have only one stereo playback system, or whose stereo signal is preferred for headphone playback.

[0048] Figure 5 is a functional block diagram of a perception model 117 of a stereo-based immersive audio encoding system according to one aspect of the disclosure, which is used to estimate direction parameters. A dominant source extraction module 1170 can extract one or more dominant sources and their directions from the M inputs. For channel-based audio inputs, source extraction or beamforming can be applied to approximately estimate one or more of the most dominant channel pairs and their directions. The directions can be interpolated between the channel pair directions of the most dominant channel pairs.

[0049] A filter bank or time-frequency frequency conversion module 1171 can convert the one or more most dominant sources from time domain to frequency domain in a plurality of subbands using a technique such as STFT. The resolution of the subbands can be determined by the characteristics of the auditory system. For example, a finer resolution at high frequencies can be selected in order to support sufficient spectral resolution to separate multiple sources in different directions. In one aspect, each subband can include a grouping of multiple frequency bins from the STFT. As mentioned, the spectral resolution used for dominant source estimation can be higher (e.g., finer) than the spectral resolution of the time-frequency tiling used for the decoded stereo signal. Since the required parameter data rate is roughly proportional to the number of subbands, the number of subbands can also depend on the target bit rate for direction parameter transmission.

[0050] The partial masking loudness module 1172 can operate on loudness estimates for subbands of the dominant source to account for masking effects when multiple competing sources partially mask each other, resulting in a dominant source with the largest loudness. The partial masking loudness module 1172 can model the masking effects by considering different spatial directions. The coded band mapping module 1173 can map the estimated loudness values in the subbands to the B subbands of the time-frequency tiles to be used at the decoder side for the stereo signal. The direction estimation module 1174 can estimate the direction of the virtual loudspeaker pair to recreate the dominant sound location in each subband as an azimuth and elevation (in degrees) relative to a default listener position.

[0051] In implementations, the intended perceptual source direction is typically known exactly only for object-based audio with corresponding metadata. In one aspect, the source extraction module 1170 is not used, and the direction estimation is based on the object signal loudness after the masking effects and the metadata. For high-fidelity ambisonics, source extraction or beamforming can be applied to approximately estimate the most dominant source and its direction.

[0052] Figure 6 is a functional block diagram of a perceptual model 117 of a stereo-based immersive audio encoding system according to one aspect of the disclosure for estimating a dominant sound and its associated virtual loudspeaker pair direction from channel-based input. As Figure 5 The filter bank or time-frequency conversion module 1171 can convert the M input sources from time domain to frequency domain in a plurality of subbands.

[0053] The loudness model 1175 can operate on loudness estimates from each input channel to model the masking effects and consider direction estimation based on the input channel layout. The loudness model 1175 can perform a triangulation between loudspeaker positions of the two or three loudest channels to account for phantom images. Thus, the direction estimation takes into account the input channel layout. Using Figure 6 The virtual loudspeaker pair direction of the dominant sound estimated using the channel-based input model of Figure 5 may be more computationally efficient than the source extraction model of The coded band mapping module 1173 can map the estimated loudness values in the subbands to the B subbands of the stereo signal at the encoder side. The direction estimation module 1176 can estimate the virtual loudspeaker pair direction in each subband as an azimuth and elevation (in degrees) relative to a default listener position based on the input channel layout.

[0054] For object-based audio, the source direction is typically given by metadata. The object metadata typically describes the object location, size, and can be used by a renderer such as Figure 4other characteristics of the desired source object image. Objects located within a segment of the sphere of the playback channel layout can be rendered as to be transmitted to the binaural signal as shown at the decoder side. However, since the object location is known, the perceptual model 117 can not need to estimate the source direction of this object. Instead, the azimuth and elevation of the virtual channel pair for which the object is rendered is used. Figure 4

[0055] Figure 7 depicts the use of virtual channel pairs for object rendering when the perceptual model 117 of a stereo-based immersive audio encoding system uses the azimuth / elevation of the virtual channel pair as metadata, according to one aspect of the disclosure. Figure 7 One virtual channel pair and two audio objects are shown, which both appear when the binaural signal rendered by the virtual channel pair is played back. Object 1 is a dry point source, which is rendered by copying the mono object signal to the right channel only. Object 2 is rendered by adding some reverb to increase the perceived distance and some decorrelation between the left and right channels, and the object is panned to the right. The downmix signal is generated by adding both rendered signals. The STIC metadata of the source direction is the azimuth / elevation of the virtual channel pair. Since the virtual channel pair angle is usually different from the source angle of the phantom image produced by the virtual channel pair, this direction is usually different from the object metadata.

[0056] Objects in the same segment of the sphere can be rendered to different virtual channel pairs to achieve better spatial resolution and optimized STIC rendering quality. When multiple virtual channel pairs are used, the perceptual model 117, such as the loudness model 1175 of Figure 6 , can estimate which virtual channel pair dominates in each time-frequency tile of the decoded binaural signal by estimating the loudness produced by each virtual channel pair after the masking effect.

[0057] For HOA-based signals, the dominant source signals and directions can be derived by a singular value decomposition (SVD) approach. Then, the perceptual model 117 can process these dominant source signals and directions in the same way as for object signals to derive the partially masked loudness.

[0058] Referring back to Figure 4 , the weighting control module 123 can generate the weighting factors w c and w d ​These weighting factors are applied to the corresponding PxBstamps of the stereo signal to generate the weighted time-frequency stamps of the P output channel pairs. The weighting control module 123 can control the spatial rendering by generating the weighting factors w c and w d for the PxBstamps based on the playback channel layout, the direction of the virtual loudspeaker pair of the dominant sound, and the results of the analysis performed by the audio analysis module 121 on the decoded stereo signal. d The output of the time-frequency stamp separator 111 is split into two paths, one of which has a decorrelator that applies the weighting factor w c to reduce the correlation between the channel pairs. The decorrelation can be applied to reduce comb filtering effects that can cause large perceived image shifts when the listener moves. The amount of decorrelation can be controlled by the ratio of the weighting factors w d and w c .

[0059] Figure 8 is a functional block diagram of the processing of a channel pair of a stereo-based immersive audio encoding system according to an aspect of the disclosure. The decoded stereo downmix signal 801 can be divided into frames and processed by the time-frequency stamp separator 111 to convert the left and right signals from the time domain to B subbands in the frequency domain. The left and right signals 803 for the B subbands are fed into P parallel processing units representing P pairs of output channels. Each processing unit can include two multiplier 830, a decorrelator 832, an adder 834, and a time-frequency stamp merger module 836. In the processing unit, the left and right signals 803 can receive the same parallel processing for the left and right channels of the pair.

[0060] The left and right signals 803 in each processing unit are split into two paths, one of which is multiplied by the weighting factor w d and the second of which as a decorrelator path is multiplied by the weighting factor w c The weighting factors w d and w c,1 for the P pairs of output channels can be indexed as {w c,2 ,w c,P ,…w d,1} and {w d,2 ,w d,P ,…w c,1 In one aspect, the same set of {w c,2 ,w c,P ,…w d,1 ,w d,2 ,…w d,P}may be applied across all B subbands of signal 803. The output in multiplier 830 for the decorrelator path is applied to decorrelator 125. Decorrelator 125 in each processing unit decorrelates the left and right signals of w d The weighted signals are filtered to remove the correlation of the corresponding channel pair with all other channel pairs, but not intended to change the correlation between the left and right channels of that channel pair. Adder 834 sums the decorrelated output 805 in decorrelator 125 with the corresponding left and right signals of unprocessed output 807 to generate a weighted output signal 809 for the channel pair. By weighting and adding the decorrelated output 805 of a channel pair with the unprocessed output 807 in adder 834, the ratio of the weighting factors w c The ratio of w c and w d controls the amount of decorrelation of the weighted output signal 809 for that channel pair.

[0061] The processing units can perform the weighted addition of the decorrelated output 805 with the unprocessed output 807 to generate a weighted output signal 809 for each of the B subbands. Time-frequency tile merger module 113 converts the weighted output signals 809 for the B subbands of each channel pair from the frequency domain back to the time domain to generate a channel pair signal 811 for each channel pair. Channel pair combiner module 131 combines the channel pair signals 811 from the P channel pairs of the output channel layout to generate audio signals 813 for the N loudspeakers of the playback system. In one aspect, N can equal 2xP, and the left and right signals of each channel pair signal 811 can drive the left and right loudspeakers of the corresponding channel pair. In one aspect, the left and right signals can be combined to drive a single loudspeaker.

[0062] For an implementation based on STFT processing, in mathematical terms, time- frequency tile separator 111 converts the left and right channel signals of stereo downmix signal 801, l mix and r mix to an STFT representation:

[0063] L mix (k) = STFT(l mix (n)) (Equation 1)

[0064] R mix (k) = STFT(r mix (n))

[0065] where n is a time-domain sample index and k is an STFT bin index.

[0066] The weighted output signal 809 for each channel pair is computed by adding the decorrelated output 805 with the unprocessed output 807:

[0067] L out (p, k) = w c (p, b) L mix (k) + Decorr(w d (p, b) L mix (k)) (Equation 2)

[0068] R out (p, k) = w c (p, b) R mix (k) + Decorr(w d (p, b) R mix (k))

[0069] where p is a channel pair index, b is a subband index, w c (p, b) is a weighting factor w c , w d (p, b) is a weighting factor w d for channel pair p and subband b. Each subband can include a grouping of STFT bins.

[0070] The time-frequency collage merger module 113 converts the complex STFT spectrum of the weighted output signal 809 back to the time domain of the channel pair signal 811:

[0071] l out (p, n) = STFT -1 (L out (p, k)) (Equation 3)

[0072] r out (p, n) = STFT -1 (R out (p, k))

[0073] The weighting factors w c and w d can be calculated by the following equations:

[0074] w Pan (p, b) = PanningWeight(a, e) (Equation 4)

[0075] w(p, b, f) = (1 - w smooth )w Pan (p, b) + w smooth w Pan (p, b, f - 1) (Equation 5)

[0076]

[0077]

[0078] where PanningWeight() is a function that computes a panning weight factor w for a channel pair p and a subband b based on the azimuth angle a and the elevation angle e given the geometry of the target channel layout Pan (p,b). In one aspect, the azimuth angle a and the elevation angle e can comprise the azimuth and elevation of a virtual loudspeaker pair to recreate a dominant source received from the perceptual model 117. For example, the left loudspeaker of the virtual pair is located at {-a, e} and the right loudspeaker is located at {a, e}. To reduce or minimize spectral distortion due to spatial rendering, a temporal smoothing of the weight factor can be performed. smooth is a smoothing factor that can depend on signal characteristics of the downmix signal 801, e.g., the prediction gain and the onset strength in a signal analysis performed by the audio analysis module 121. In one aspect, w smooth may be the same for all P channel pairs and B subbands. The weight factor w corr controls the amount of decorrelation applied by controlling the ratio between w c (p,b) and w d (p,b). The weight factor w corr may also depend on the prediction gain and the onset strength of the downmix signal 801. In one aspect, w corr may be the same for all P channel pairs and B subbands. The frame index f indicates the current STFT frame. A smoothing of w(p,b,f) can be performed for subsequent frames. In one aspect, w Pan (p,b), w(p,b,f), w c (p,b) and w d (p,b) can be independent of the subband.

[0079] Figure 9 is a functional block diagram of an audio analysis module 121 of a stereo-based immersive audio encoding system for adjusting a weight factor according to one aspect of the present disclosure. Each channel of a decoded stereo signal, such as a stereo downmix signal 801, can be processed in the time domain by a forward predictor 1211. The forward predictor 1211 can generate a prediction signal 901 that is subtracted from the actual decoded stereo signal to generate a prediction error signal 903. A prediction gain estimator 1212 can estimate a prediction gain based on an estimated difference of the RMS levels of the decoded stereo signal and the prediction error signal 903. In parallel, an onset / instantaneous detector 1213 evaluates the envelope of the decoded stereo signal to estimate the strength of the onset. The maximum of the results in both channels is used for further processing.

[0080] The prediction gain is an indication of the temporal "smoothness" of the decoded audio signal. For audio signals with high prediction gain, the weight factor can need more smoothing. The weight factor wc and w d is time-smoothed, and more decorrelation can be applied. On the other hand, if the onset strength is significant, the weighting factor w c and w d can be reduced, while less decorrelation can be applied. If the onset strength is high, the time-frequency tiled audio signal can be limited to mainly a single playback channel pair to avoid time smearing and spectral distortion. Thus, the weighting factor w c and w d can be limited such that only one channel pair carries most of the signal energy, while all other channel pairs have negligible energy. In one aspect, the encoder side can perform a signal analysis on the stereo signal to estimate its onset strength and prediction gain. The encoder side can transmit parameters corresponding to the onset strength and prediction gain of the encoded stereo signal to the decoder for use as described.

[0081] Figure 10 is a functional block diagram of a weighting control module 123 for generating weighting factors for time-frequency tiling according to one aspect of the disclosure. A first estimator module 1231 can estimate a temporal fluctuation of the direction parameters for time-frequency tiling. A second estimator module 1232 can compute an initial estimate of the parameters for time-smoothing of the weighting factors based on the temporal fluctuation of the direction parameters estimated in the first estimator module 1231, such as the smoothing factor w smooth in Equation 5. A weighting factor generation module 1233 can generate the weighting factors, such as w(p, b, f) of Equation 6 for P channel pairs and B subbands of frame f, based on the initial estimate of the time-smoothing parameters, the azimuth angle a and the elevation angle e of the virtual loudspeaker pair for the subband received through the direction parameters, the prediction gain and the onset strength from the audio analysis module 121, and the playback channel layout.

[0082] A decorrelation estimator module 1234 can control the amount of decorrelation applied by generating the weighting coefficients w corr of Equations 6 and 7 based on the prediction gain and the onset strength as described, thereby controlling the amount of decorrelation applied. As mentioned, decorrelation can be applied to avoid comb-filter effects that can cause large image shifts when the listener moves. These effects are most noticeable in signals with smooth envelopes and high prediction gain. However, when decorrelation is applied, it can also cause an increase in audible reverberation, and the signal source can appear more distant compared to the input signal.

[0083] Due to the perceived distance and reverberation modification, the use of decorrelation is reduced or minimized and applied only when necessary. This can be achieved by the decorrelation estimator module 1234, specifically controlling the decorrelation through the generation of the weighting coefficients w corr . The weighting coefficients w corrw(p, b, f) applicable in the weighting factor generation module 1233 to generate w of formulas 6 and 7 c (p, b) and w d (p, b). The weighting factor w c (p, b) and w d (p, b) can be used to adaptively process the time-frequency tiles to reduce or minimize spectral distortion due to spatial positioning.

[0084] Since the weighting factor w d is applied to the time-frequency tiles before the decorrelator 125 instead of after, only those parts of the decoded stereo signal that need to be decorrelated enter the decorrelation 125. If the weighting factor w d is applied after instead of before the decorrelator 125, large onsets that do not need to be decorrelated can temporarily spread into parts of the decoded stereo signal that need to be decorrelated and thus can cause a reverberation artifact. In addition, by excluding the output channel pair with the largest energy from the decorrelator processing, the use of the decorrelator 125 can be reduced or minimized in each time-frequency tile. This is possible because this channel pair is not related to any other channel pair processed by the decorrelator 125.

[0085] The weighting factors can be balanced such that the input signal loudness is preserved. In one aspect, as a first approximation, the RMS value of the weighting factors for all P channel pairs in a time-frequency tile can be set to 1. By normalizing using a frequency-dependent exponent σ between 1.0 and 2.0 (with smaller values at lower frequencies), a more accurate loudness match can be achieved and coloration can be prevented:

[0086]

[0087] where w c (p) and w d (p) is w c (p, b) and w d (p, b).

[0088] Figure 4 The stereo-based immersive audio encoding system is based on a single stereo downmix of the audio content. This means, for example, that any back channel content can be mixed with front channel content, which in turn can also lead to different positioning after spatial rendering if the signals overlap in time and frequency. To improve the positioning accuracy, multiple downmixes can be used, where each downmix only includes those signals that are located in a sector of a sphere represented by the downmix. All sectors can cover the entire sphere without overlap.

[0089] Figure 11Downmixes of audio channels for multiple sectors for seven loudspeaker layouts are depicted in accordance with one aspect of the disclosure. Figure 11 An example of generating two downmixes, one for channels in the front sector of a 7.0 layout, and one for channels in the back sector, is shown. For example, for a layout with a sky channel (such as 7.0.4), the sky channel can be assigned to a sector using the same mapping.

[0090] Figure 12 is a functional block diagram of a stereo-based immersive audio coding system that encodes and decodes multiple sections or sectors of a loudspeaker layout in accordance with one aspect of the disclosure. A section separation module 133 can separate a spherical surface of a channel layout into multiple sections or sectors. Figure 1 Multiple instances of the STIC system ofare used to encode signals associated with each section of the spherical surface. At the decoder side, the audio output signals in each section are added to generate the final audio output for the playback system. In one aspect, Figure 4 Multiple instances of the STIC system of may be used to encode and decode signals associated with multiple sections. In general, the sections can have any number and any shape. However, for channel-based audio, the sections are typically symmetric across the median plane. To achieve a good tradeoff of bit rate versus quality, the number of sections should be as small as possible, but large enough to achieve the required localization accuracy.

[0091] In one aspect of a hybrid stereo-based immersive audio coding system, it can be advantageous to remove one channel (such as the front center channel) from the remaining channels when applying the STIC technique. The front center channel can be encoded, decoded independently of the STIC system, and added to the remaining channels that are rendered using the STIC system. This hybrid configuration can improve the rendered image of the front center channel, which is typically used for dialogue in movies and television content. Figure 4

[0092] Figure 13 is a functional block diagram of a hybrid stereo-based immersive audio coding system that encodes and decodes a single channel, such as a center channel, independently of other channels that are encoded and decoded using a STIC system in accordance with one aspect of the disclosure. In one example, the input channels of a surround signal can have a 5.1 layout, including 2 channel pairs (a left-right channel pair, a left surround and right surround channel pair), and two single channels (center and LFE).

[0093] A channel pair extraction module 141 can extract all channel pairs (such as the left-right channel pair and the left surround and right surround channel pair) for encoding and decoding by Figure 1 , Figure 4 or Figure 12The STIC system. The mono channel extraction module 143 can extract mono channels (such as center and LFE) to be encoded independently of the STIC system. In one aspect, the audio codec 145 can encode the extracted mono channels. Information about the presence and location of the mono channels can be added to the STIC parameters so that the decoder can correctly render the channels.

[0094] At the decoder side, the decoder of the audio codec 145 can decode the mono channels. The mono channel renderer 147 can render the decoded mono channels to the output layout indicated by the playback channel layout. For example, if the output layout has a speaker location at a mono channel location (such as a front center speaker), the decoded mono channel for the center channel can be passed to the front center speaker. Otherwise, the decoded mono channel for the center channel can be rendered to the closest available channel. In one aspect, virtual sound source positioning techniques (such as vector-based amplitude panning (VBAP)) can be used.

[0095] The channel combiner module 149 can add the rendered mono channels to the channel pairs rendered by the STIC system to generate the reconstructed audio signal. For example, if the playback channel layout has a front center channel, the channel combiner module 149 can route the signal rendered for the mono center channel to the front center channel, or the channel combiner module 149 can add the signal for the mono center channel rendered to a channel pair to the corresponding channel pair signal rendered by the STIC system. In one aspect, if there is an LFE channel, the mono channel for the LFE can be routed to the LFE channel of the playback channel layout.

[0096] Figure 14 is a flow diagram of an encoder-side processing method 1400 of a stereo-based immersive audio encoding system to generate stereo signals and direction parameters from an immersive audio signal according to an aspect of the disclosure. The method 1400 can be implemented by an encoder-side of a STIC system of Figure 1 , Figure 4 , Figure 12 or Figure 13 .

[0097] In operation 1401, the method 1400 generates a two-channel stereo signal from an immersive audio signal. The immersive audio signal can include a plurality of audio channels of an input channel layout, a plurality of audio objects, or HOA. In one aspect, a downmixer module can downmix the multi-channel input to the stereo signal, or a renderer module can render the plurality of audio objects or HOA to the stereo signal.

[0098] In operation 1403, the method 1400 generates direction parameters based on the audio content, the direction parameters describing optimal virtual loudspeaker pair directions to recreate a perceived dominant sound location of the audio content in a plurality of frequency subbands. The virtual loudspeaker pair directions for each of the subbands can be given as an azimuth and an elevation (in degrees) relative to a default listener position.

[0099] In operation 1405, the method 1400 transmits the binaural stereo signal and the direction parameters to a decoding device over a communication channel. The communication channel can have a limited bandwidth. The bandwidth requirement of the direction parameters can be significantly lower than the bandwidth requirement of the individual audio channels of the stereo signal.

[0100] Figure 15 is a flowchart of a decoder-side processing method 1500 of a binaural-based immersive audio encoding system to reconstruct an immersive audio signal for a multi-channel playback system according to an aspect of the disclosure. The method 1500 can be implemented by a decoder-side of a STIC system of Figure 1 、 Figure 4 、 Figure 12 or Figure 13 .

[0101] In operation 1501, the method 1500 receives a binaural stereo signal and direction parameters from an encoding device, the direction parameters describing optimal virtual loudspeaker pair directions to recreate a perceived dominant sound location of an audio content represented by the binaural stereo signal in a plurality of frequency subbands. The audio content can be a multi-channel immersive audio signal.

[0102] In operation 1503, the method 1500 generates a plurality of time-frequency tiles for a plurality of channel pairs of the playback system from the binaural stereo signal, the plurality of time-frequency tiles representing a frequency-domain representation of each channel of the binaural stereo signal in the plurality of frequency subbands. The number of subbands B can be determined by a required spectral resolution. The binaural stereo signal can be divided into frames represented by the time-frequency tiles. The frequency-domain representation of the stereo signal can be separated or duplicated into P parallel processing paths, where each processing path can be associated with each channel pair of the playback system.

[0103] In operation 1505, the method 1500 generates weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs based on the direction parameters. In one aspect, the weighting factors can be generated based on the virtual loudspeaker pair directions to recreate a perceived dominant sound location of the audio content represented by the binaural stereo signal in the plurality of frequency subbands, an analysis of the stereo signal, and an output channel layout of the playback system. In one aspect, the weighting factors can be controlled to reduce correlation between the channel pairs.

[0104] In operation 1507, the method 1500 applies a plurality of weighting factors to the plurality of time-frequency tiles to spatially render the time-frequency tiles through a plurality of channels of a playback system. These weighting factors can be used to adaptively process the time-frequency tiles (such as using a decorrelator) to reduce or minimize spectral distortion due to spatial rendering.

[0105] Embodiments of the stereo-based immersive audio encoding techniques described herein can be implemented in a data processing system, such as a network computer, a network server, a tablet computer, a smart phone, a laptop computer, a desktop computer, other consumer electronics device, or other data processing system. In particular, the operations described for the stereo-based immersive encoding system are digital signal processing operations performed by a processor executing instructions stored in one or more memories. The processor can read the stored instructions from the memories and execute the instructions to perform the described operations. These memories represent examples of machine-readable non-transitory storage media that can store or contain computer program instructions that, when executed, cause the data processing system to perform one or more of the methods described herein. The processor can be a processor in a local device such as a smart phone, a processor in a remote server, or a distributed processing system of multiple processors in a local device and a remote server, with their respective memories containing portions of the instructions needed to perform the described operations.

[0106] While certain example embodiments have been described herein, it is to be understood that the example embodiments described herein are not intended to limit the broad disclosure, which can be embodied in various forms. Therefore, the description and drawings are to be regarded as illustrative in nature and not restrictive.

Claims

1. A method of encoding audio content, the method comprising: generating, by an encoding device, a binaural signal from the audio content; generating, by the encoding device, one or more direction parameters based on the audio content, each of the one or more direction parameters describing a direction of a respective pair of virtual loudspeakers that together recreate, in a respective frequency subband of a plurality of frequency subbands, a perceived dominant sound location of the audio content, wherein the respective pair of virtual loudspeakers is a pair of virtual loudspeakers positioned symmetrically left and right relative to a listener facing forward axis; and communicating, to a decoder, the binaural signal and the direction parameter over a communication channel or over a storage device.

2. The method of claim 1, wherein the audio content comprises one or more of a multichannel signal associated with a loudspeaker layout, a plurality of audio objects, or an arbitrary order Ambisonics.

3. The method of claim 1, wherein generating the direction parameter comprises: converting, by the encoding device, the audio content provided by a multichannel signal associated with a loudspeaker layout to a plurality of subbands of a frequency domain representation of the audio content; and determining, by the encoding device, a maximum loudness of the audio content for each subband of the plurality of subbands using a loudness masking model based on the loudspeaker layout associated with the multichannel signal, wherein the perceived dominant sound location is a location of the maximum loudness.

4. The method of claim 1, wherein each of the direction parameters comprises an azimuth angle and an elevation angle relative to a default listener position.

5. The method of claim 1, wherein generating the direction parameter comprises: rendering, by the encoding device, the audio content provided by a plurality of audio objects to one or more pairs of virtual channels to create an image of the plurality of audio objects; and determining, by the encoding device, a maximum loudness of the image of the plurality of audio objects created by the one or more pairs of virtual channels, wherein the perceived dominant sound location is a location of the maximum loudness.

6. The method of claim 1, further comprising: dividing the audio content into a plurality of segments based on a layout of a plurality of audio sources providing the audio content, wherein generating the binaural signal from the audio content comprises: generating a plurality of binaural signals respectively corresponding to the audio content in the plurality of segments; wherein generating the direction parameter comprises: generating a plurality of direction parameters respectively corresponding to the audio content in the plurality of segments, each of the plurality of direction parameters describing a direction of a pair of virtual loudspeakers to recreate, in a plurality of frequency subbands, a perceived dominant sound location of the audio content in a corresponding segment of the plurality of segments, and wherein communicating the binaural signal and the direction parameter comprises: communicating, to the decoder, the plurality of binaural signals and the plurality of direction parameters over the communication channel or over the storage device. ​ 7. The method of claim 1, further comprising: analyzing the binaural stereo signal to generate content analysis parameters; and communicating the content analysis parameters to the decoder.

8. The method of claim 7, wherein the content analysis parameters include parameters representing a prediction gain and a sound onset strength of the stereo signal.

9. A system configured to encode audio content, the system comprising: a memory configured to store instructions; a processor coupled to the memory and configured to execute the instructions stored in the memory to: generate a binaural stereo signal from the audio content; generate one or more direction parameters based on the audio content, each of the one or more direction parameters describing a direction of a respective pair of virtual loudspeakers that together recreate a perceived dominant sound location of the audio content in a respective frequency subband of a plurality of frequency subbands, wherein the respective pair of virtual loudspeakers is a pair of virtual loudspeakers positioned symmetrically left and right relative to a listener facing forward axis; and communicate the binaural stereo signal and the direction parameters to a decoder over a communication channel or over a storage device.

10. The system of claim 9, wherein the audio content comprises one or more of a multi-channel signal associated with a loudspeaker layout, a plurality of audio objects, or an arbitrary order Ambisonics.

11. The system of claim 9, wherein to generate the direction parameters, the processor further executes the instructions stored in the memory to: convert the audio content provided by a multi-channel signal associated with a loudspeaker layout to a plurality of subbands of a frequency domain representation of the audio content; and determine a maximum loudness of the audio content for each subband of the plurality of subbands using a loudness masking model based on the loudspeaker layout associated with the multi-channel signal, wherein the perceived dominant sound location is a location of the maximum loudness.

12. The system of claim 9, wherein each of the direction parameters comprises an azimuth angle and an elevation angle relative to a default listener position.

13. The system of claim 9, wherein to generate the direction parameters, the processor further executes the instructions stored in the memory to: render the audio content provided by a plurality of audio objects to one or more pairs of virtual channels to create an image of the plurality of audio objects; and determine a maximum loudness of the image of the plurality of audio objects created by the one or more pairs of virtual channels, wherein the perceived dominant sound location is a location of the maximum loudness.

14. The system of claim 9, wherein the processor further executes the instructions stored in the memory to: divide the audio content into a plurality of segments based on a layout of a plurality of audio sources providing the audio content, wherein the binaural stereo signal is to be generated from the audio content, the processor further executes the instructions stored in the memory to: generate a plurality of binaural stereo signals respectively corresponding to the audio content in the plurality of segments; wherein the direction parameters are to be generated, the processor further executes the instructions stored in the memory to: generate a plurality of direction parameters respectively corresponding to the audio content in the plurality of segments, each of the plurality of direction parameters describing a direction of a pair of virtual loudspeakers to recreate a perceived dominant sound location of the audio content in a corresponding segment of the plurality of segments in a plurality of frequency subbands, and wherein the binaural stereo signal and the direction parameters are to be transmitted, the processor further executes the instructions stored in the memory to: transmit the plurality of binaural stereo signals and the plurality of direction parameters to the decoder through the communication channel or through the storage device.

15. The system of claim 9, wherein the processor further executes the instructions stored in the memory to: analyze the binaural stereo signal to generate content analysis parameters; and transmit the content analysis parameters to the decoder.

16. The system of claim 15, wherein the content analysis parameters comprise parameters representative of a prediction gain and a sound onset strength of the stereo signal.

17. A method of decoding audio content, the method comprising: receiving, by a decoder device from an encoding device, a binaural stereo signal and one or more direction parameters, each of the one or more direction parameters describing a direction of a respective pair of virtual loudspeakers, wherein one or more of the respective pair of virtual loudspeakers recreate together a perceived dominant sound location of the audio content represented by the binaural stereo signal in a respective frequency subband of a plurality of frequency subbands, wherein the respective pair of virtual loudspeakers are a pair of virtual loudspeakers positioned left-right symmetrically with respect to a forward-facing axis of a listener; generating, by the decoder device from the binaural stereo signal, a plurality of time- frequency tiles of a plurality of channel pairs of a playback system, the plurality of time- frequency tiles representing a frequency domain representation of each channel of the binaural stereo signal in the plurality of frequency subbands; generating a plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs based on the direction parameters; and applying the plurality of weighting factors to the plurality of time-frequency tiles to spatially render the time-frequency tiles by the plurality of channel pairs of the playback system.

18. The method of claim 17, wherein applying the plurality of weighting factors to the plurality of time-frequency tiles comprises: applying the plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs to two channels of a corresponding one of the plurality of time- frequency tiles and the plurality of channel pairs to recreate the perceived dominant sound direction of the audio content for the plurality of frequency subbands by the plurality of channel pairs of the playback system.

19. The method of claim 17, wherein the plurality of time-frequency tiles are generated by the decoder device from the binaural stereo signal by: performing a time-frequency transform on the binaural stereo signal to generate a plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands; generating a plurality of time-frequency tiles of a plurality of channel pairs of a playback system from the plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands, each of the plurality of time-frequency tiles representing a frequency domain representation of each channel of the binaural stereo signal in a corresponding one of the plurality of frequency subbands; and generating a plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs based on the direction parameters.

20. The method of claim 17, wherein the plurality of time-frequency tiles are generated by the decoder device from the binaural stereo signal by: performing a time-frequency transform on the binaural stereo signal to generate a plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands; generating a plurality of time-frequency tiles of a plurality of channel pairs of a playback system from the plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands, each of the plurality of time-frequency tiles representing a frequency domain representation of each channel of the binaural stereo signal in a corresponding one of the plurality of frequency subbands; and generating a plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs based on the direction parameters.

21. The method of claim 17, wherein the plurality of time-frequency tiles are generated by the decoder device from the binaural stereo signal by: performing a time-frequency transform on the binaural stereo signal to generate a plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands; generating a plurality of time-frequency tiles of a plurality of channel pairs of a playback system from the plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands, each of the plurality of time-frequency tiles representing a frequency domain representation of each channel of the binaural stereo signal in a corresponding one of the plurality of frequency subbands; and generating a plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs based on the direction parameters.

22. The method of claim 17, wherein the plurality of time-frequency tiles are generated by the decoder device from the binaural stereo signal by: performing a time-frequency transform on the binaural stereo signal to generate a plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands; generating a plurality of time-frequency tiles of a plurality of channel pairs of a playback system from the plurality of frequency domain representations of the binaural stereo signal in the plurality of frequency subbands, each of the plurality of time-frequency tiles representing a frequency domain representation of each channel of the binaural stereo signal in a corresponding one of the plurality of frequency subbands; and generating a plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs based on the direction parameters.

19. The method of claim 17, wherein the plurality of weighting factors comprises a plurality of decorrelation weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs, and wherein applying the plurality of weighting factors to the plurality of time-frequency tiles comprises: applying the plurality of decorrelation weighting factors for the plurality of time- frequency tiles of the plurality of channel pairs to corresponding ones of the plurality of time- frequency tiles and the plurality of channel pairs to reduce correlation between the plurality of channel pairs.

20. The method of claim 17, wherein generating the plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs comprises: generating a characteristic of the binaural stereo signal; and generating the plurality of weighting factors based on the characteristic of the binaural stereo signal, a layout of the plurality of channel pairs of the playback system, and the direction parameter.

21. The method of claim 20, wherein generating a characteristic of the binaural stereo signal comprises: analyzing the binaural stereo signal to generate a prediction gain based on a forward prediction of the binaural stereo signal, wherein the prediction gain measures a temporal smoothness of the binaural stereo signal; and analyzing the binaural stereo signal to generate a onset strength, wherein the onset strength estimates an onset strength of the binaural stereo signal.

22. The method of claim 21, wherein generating the plurality of weighting factors based on the characteristic of the binaural stereo signal comprises: controlling the weighting factors for the plurality of time-frequency tiles to cause one of the channel pairs to carry a majority of signal energy of the binaural stereo signal when the onset strength is strong.

23. The method of claim 21, wherein generating the plurality of weighting factors based on the characteristic of the binaural stereo signal comprises: generating a plurality of decorrelation weighting factors for the plurality of time- frequency tiles of the plurality of channel pairs based on the prediction gain and the onset strength, wherein the plurality of decorrelation weighting factors are applied to the plurality of time- frequency tiles of the plurality of channel pairs to reduce correlation between the plurality of channel pairs.

24. The method of claim 20, wherein generating the plurality of weighting factors based on the characteristic of the binaural stereo signal, the layout of the plurality of channel pairs of the playback system, and the direction parameter comprises: estimating a temporal fluctuation of the direction parameter in the plurality of frequency subbands; and determining a smoothing factor to temporally smooth the plurality of weighting factors based on the estimated temporal fluctuation of the direction parameter.

25. The method of claim 20, wherein generating the plurality of weighting factors based on the characteristic of the binaural stereo signal, the layout of the plurality of channel pairs of the playback system, and the direction parameter comprises: controlling the plurality of weighting factors for the plurality of channel pairs to distribute signal energy of the binaural stereo signal across the plurality of channel pairs to spatially localize a perceived image of the audio content.

26. A system configured to decode audio content, the system comprising: a memory configured to store instructions; a processor coupled to the memory and configured to execute the instructions stored in the memory to: receive, from an encoding device, a binaural stereo signal and one or more direction parameters, each of the one or more direction parameters describing a direction of a respective pair of virtual loudspeakers, wherein one or more of the respective pair of virtual loudspeakers together recreate, in a respective frequency subband of a plurality of frequency subbands, a perceived dominant sound location of the audio content represented by the binaural stereo signal, wherein the respective pair of virtual loudspeakers are a pair of virtual loudspeakers positioned symmetrically left and right relative to a forward-facing axis of a listener; generate, from the binaural stereo signal, a plurality of time-frequency tiles of a plurality of channel pairs of a playback system, the plurality of time-frequency tiles representing a frequency-domain representation of each channel of the binaural stereo signal in the plurality of frequency subbands; generate, based on the direction parameters, a plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs; and apply the plurality of weighting factors to the plurality of time-frequency tiles to spatially render the time-frequency tiles by the plurality of channel pairs of the playback system.

27. The system of claim 26, wherein to apply the plurality of weighting factors to the plurality of time-frequency tiles, the processor further executes the instructions stored in the memory to: apply the plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs to both channels of a corresponding one of the plurality of time-frequency tiles and the plurality of channel pairs to recreate, by the plurality of channel pairs of the playback system, the perceived dominant sound direction of the audio content for the plurality of frequency subbands.

28. The system of claim 26, wherein the plurality of weighting factors comprise a plurality of decorrelation weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs, and wherein to apply the plurality of weighting factors to the plurality of time-frequency tiles, the processor further executes the instructions stored in the memory to: apply the plurality of decorrelation weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs to the corresponding one of the plurality of time-frequency tiles and the plurality of channel pairs to reduce correlation between the plurality of channel pairs.

29. The system of claim 26, wherein to generate the plurality of weighting factors for the plurality of time-frequency tiles of the plurality of channel pairs, the processor further executes the instructions stored in the memory to: generate a characteristic of the binaural stereo signal; and generate the plurality of weighting factors based on the characteristic of the binaural stereo signal, a layout of the plurality of channel pairs of the playback system, and the direction parameters.

30. The system of claim 29, wherein to generate a characteristic of the binaural stereo signal, the processor further executes the instructions stored in the memory to: generate a characteristic of the binaural stereo signal; and generate the plurality of weighting factors based on the characteristic of the binaural stereo signal, a layout of the plurality of channel pairs of the playback system, and the direction parameters. analyzing the binaural signal to generate a prediction gain based on a forward prediction of the binaural signal, wherein the prediction gain measures a temporal smoothness of the binaural signal; and analyzing the binaural signal to generate a onset strength, wherein the onset strength estimates an onset strength of the binaural signal.

31. The system of claim 30, wherein the plurality of weighting factors are to be generated based on the characteristics of the binaural signal, the processor further executes the instructions stored in the memory to: control the weighting factors for the plurality of time-frequency tiles to cause one of the pairs of channels to carry a majority of signal energy of the binaural signal when the onset strength is strong.

32. The system of claim 30, wherein the plurality of weighting factors are to be generated based on the characteristics of the binaural signal, the processor further executes the instructions stored in the memory to: generate a plurality of decorrelation weighting factors for the plurality of time- frequency tiles of the plurality of pairs of channels based on the prediction gain and the onset strength, wherein the plurality of decorrelation weighting factors are applied to the plurality of time-frequency tiles of the plurality of pairs of channels to reduce correlation between the plurality of pairs of channels.

33. The system of claim 29, wherein the plurality of weighting factors are to be generated based on the characteristics of the binaural signal, the layout of the plurality of pairs of channels of the playback system, and the direction parameter, the processor further executes the instructions stored in the memory to: estimate a temporal fluctuation of the direction parameter in the plurality of frequency subbands; and determine a smoothing factor to temporally smooth the plurality of weighting factors based on the estimated temporal fluctuation of the direction parameter.

34. The system of claim 29, wherein the plurality of weighting factors are to be generated based on the characteristics of the binaural signal, the layout of the plurality of pairs of channels of the playback system, and the direction parameter, the processor further executes the instructions stored in the memory to: control the plurality of weighting factors for the plurality of pairs of channels to distribute signal energy of the binaural signal across the plurality of pairs of channels to spatially localize a perceived image of the audio content.

Citation Information

Patent Citations

  • Method and apparatus for enhancement of audio reconstruction

    CN101658052A

  • Spatial audio coding based on universal spatial cues

    US20070269063A1