Head tracking for parameterized binaural output systems and methods
By analyzing the audio content in the encoder to determine the dominant audio component and encode it, and rendering the binaural output in the decoder, the high computational complexity and delay problems in the existing technology are solved, and low-complexity head tracking binaural output is achieved, which is suitable for mobile devices and loudspeakers.
Patent Information
- Application Number
- CN202110229741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-12-14
- Filing Date
- 2016-11-17
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2036-11-17
AI Technical Summary
Existing technologies have high computational complexity and latency issues in reproducing audio content, and it is particularly difficult to achieve efficient binaural output on head tracking and mobile devices.
By analyzing the audio content in the encoder, determining the dominant audio component and direction, and encoding it into a coded signal, the decoder renders the dominant component and residual component based on head tracking data to achieve low-complexity binaural output.
It achieves low-complexity binaural output under head tracking, reduces latency, is suitable for mobile devices and loudspeakers, and is compatible with matrix coding systems.
Smart Images

Figure CN113038354B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application number 201680075037.8, application date November 17, 2016, and invention name “Head tracking for parameterized binaural output system and method”. Technical Field
[0002] The present invention provides systems and methods for an improved form of parameterized binaural output, optionally utilizing head tracking.
[0003] References
[0004] Gundry, K., “A New Matrix Decoder for Surround Sound,” AES 19th International Conf., Schloss Elmau, Germany, 2001.
[0005] Vinton, M., McGrath, D., Robinson, C., Brown, P., “Next generation surround decoding and up-mixing for consumer and professional applications,” AES 57th International Conf, Hollywood, CA, USA, 2015.
[0006] Wightman, FL and Kistler, DJ (1989). “Headphone simulation of free-field listening. I. Stimulus synthesis,” J. Acoust. Soc. Am. 85, 858–867.
[0007] ISO / IEC 14496-3:2009 – Information technology — Coding of audio-visual objects — Part 3: Audio, 2009.
[0008] Mania, Katerina et al. "Perceptual sensitivity to head tracking latency in virtual environments with varying degrees of scene complexity." Proceedings of the 1st Symposium on Applied perception in graphics and visualization. ACM, 2004.
[0009] Allison, R.S., Harris, L.R., Jenkin, M., Jasiobedzka, U., and Zacher, J.E. (2001, March). Tolerance of temporal delay in virtual environments. In Virtual Reality, 2001. Proceedings. IEEE (pp. 247-254). IEEE.
[0010] Van de Par, Steven, and Armin Kohlrausch. "Sensitivity to auditory-visual asynchrony and to jitter in auditory-visual timing." Electronic Imaging. International Society for Optics and Photonics, 2000. Background Art
[0011] Any discussion of the background art throughout the specification should in no way be considered as an admission that such art is widely known or forms part of the common general knowledge in the field.
[0012] The creation, encoding, distribution, and reproduction of audio content has traditionally been channel-based. That is, a specific target playback system is envisioned for the content of the entire content ecosystem. Examples of such target playback systems are mono, stereo, 5.1, 7.1, 7.1.4, and so on.
[0013] If the content is to be reproduced on a playback system that is different from the intended playback system, downmixing or upmixing may be applied. For example, 5.1 content can be reproduced on a stereo playback system by adopting specific known downmixing equations. Another example is the playback of stereo content on a 7.1 speaker setup, which may include a so-called upmixing process that may or may not be guided by information present in the stereo signal, such as the stereo signal used by a so-called matrix encoder such as Dolby Pro Logic. In order to guide the upmixing process, information about the original position of the signal before downmixing may be implicitly signaled by including specific phase relationships in the downmix equations, or in other words, by applying complex-valued downmix equations. A well-known example of such a downmixing method that uses complex-valued downmix coefficients for content with speakers placed in two dimensions is LtRt (Vinton et al., 2015).
[0014] The resulting (stereo) downmix signal can be reproduced on a stereo loudspeaker system, or can be upmixed to a loudspeaker setup with surround speakers and / or height speakers. The desired positioning of the signal can be derived by the upmixer from the inter-channel phase relationship. For example, in an LtRt stereo representation, a signal that is out of phase (e.g., having an inter-channel waveform normalized cross-correlation coefficient close to -1) should ideally be reproduced by one or more surround speakers, while a positive correlation coefficient (close to +1) indicates that the signal should be reproduced by the speakers in front of the listener.
[0015] Various upmixing algorithms and strategies have been developed, differing in their strategies for recreating a multichannel signal from a stereo downmix. In relatively simple upmixers, the normalized cross-correlation coefficients of the stereo waveform signals are tracked over time, and the signal(s) are steered to either the front or rear speakers based on the value of the normalized cross-correlation coefficients. This approach works well for relatively simple content, where only a single auditory object is present simultaneously. More advanced upmixers control the signal flow from the stereo input to the multichannel output based on statistical information derived from specific frequency regions (Gundry 2001, Vinton et al. 2015). Specifically, a signal model based on a leading or dominant component and a stereo (diffuse) residual signal can be employed within each time / frequency tile. In addition to estimating the dominant component and the residual signal, a directional angle (in azimuth, possibly supplemented with elevation) is also estimated, and the dominant component signal is then steered to one or more loudspeakers to reconstruct the (estimated) position during playback.
[0016] The use of matrix encoders and decoders / upmixers is not limited to channel-based content. Recent developments in the audio industry are based on audio objects, rather than channels, where one or more objects contain an audio signal and associated metadata that indicates, among other things, the expected position of the audio signal over time. For such object-based audio content, matrix encoders can also be used, as outlined in Vinton et al. (2015). In such systems, the object signal is downmixed to a stereo signal representation with downmix coefficients that depend on the object position metadata.
[0017] Upmixing and reproduction of matrix-coded content is not necessarily limited to playback on loudspeakers. A leading component or a representation of the leading component containing the dominant component signal and the (expected) position enables reproduction on headphones by means of convolution with the head-related impulse response (HRIR) (Wightman et al., 1989). Figure 1 A simple schematic diagram of a system implementing this method is shown in FIG. An input signal 2 in matrix-coded format is first analyzed 3 to determine the dominant component direction and amplitude. The dominant component signal is convolved 4, 5 with a pair of HRIRs derived from a lookup table 6 based on the dominant component direction to compute an output signal for headphone playback 7, such that the playback signal is perceived as coming from the direction determined by the dominant component analysis stage 3. This approach can be applied to wideband signals as well as to individual subbands and can be supplemented with various forms of dedicated processing of the residual (or diffuse) signal.
[0018] The use of a matrix encoder is well suited for distribution to and reproduction on AV receivers, but can be problematic for mobile applications requiring low transmission data rates and low power consumption.
[0019] Whether using channel-based or object-based content, matrix encoders and decoders rely on fairly accurate inter-channel phase relationships of the signals delivered from the matrix encoder to the decoder. In other words, the delivery format should be largely waveform-preserving. This reliance on waveform preservation can be problematic under bit-rate-constrained conditions, where audio codecs employ parametric approaches rather than waveform coding tools to achieve better audio quality. Examples of such parameterized tools that are generally known to not preserve waveforms, as implemented in the MPEG-4 Audio codec (ISO / IEC 14496-3:2009), are generally referred to as spectral band replication, parametric stereo, spatial audio coding, and the like.
[0020] As outlined in the previous section, the upmixer involves analysis and steering of the signal (or HRIR convolution). For powered devices, such as AV receivers, this generally does not cause a problem, but for battery-operated devices, such as mobile phones and tablets, the computational complexity and corresponding memory requirements associated with these processes are generally undesirable because they negatively impact battery life.
[0021] The aforementioned analysis also typically introduces additional audio latency. Such audio latency is undesirable because (1) it requires video latency to maintain audio-video lip synchronization, which requires significant memory and processing power, and (2) in the case of head tracking, it may cause asynchrony / delay between head movement and audio rendering.
[0022] Matrix-coded downmixes may also not sound optimal over stereo speakers or headphones due to the possible presence of strong out-of-phase signal components. Summary of the Invention
[0023] It is an object of the present invention to provide an improved form of parametric binaural output.
[0024] According to a first aspect of the present invention, a method for encoding channel-based or object-based input audio for playback is provided, the method comprising the following steps: (a) first rendering the channel-based or object-based input audio into an initial output representation (e.g., an initial output representation); (b) determining an estimate of a dominant audio component from the channel-based or object-based input audio, and determining a series of dominant audio component weighting factors for mapping the initial output representation to the dominant audio component; (c) determining an estimate of a dominant audio component direction or position; and (d) encoding the initial output representation, the dominant audio component weighting factors, and the dominant audio component direction or position into an encoded signal for playback. Providing a series of dominant audio component weighting factors for mapping the initial output representation to the dominant audio component can enable the determination of an estimate of the dominant component using the dominant audio component weighting factors and the initial output representation.
[0025] In some embodiments, the method further includes determining an estimate of a residual mix, the residual mix being a rendering of the initial output representation minus a dominant audio component or an estimate thereof. The method may also include generating an anechoic binaural mix of the input audio, either channel-based or object-based, and determining the estimate of the residual mix, wherein the estimate of the residual mix may be a rendering of the anechoic binaural mix minus a dominant audio component or an estimate thereof. Furthermore, the method may include determining a series of residual matrix coefficients for mapping the initial output representation to the estimate of the residual mix.
[0026] The initial output representation may include a headphone or loudspeaker representation. The channel-based or object-based input audio may be sliced in time and frequency, and the encoding steps may be repeated for a series of time steps and a series of frequency bands. The initial output representation may include a stereo speaker mix.
[0027] According to another aspect of the present invention, there is provided a method for decoding an encoded audio signal comprising: a first (e.g., initial) output representation (e.g., first / initial output representation), a dominant audio component direction, and a dominant audio component weighting factor; the method comprising the steps of: (a) determining an estimated dominant component using the dominant audio component weighting factor and the initial output representation; (b) rendering the estimated dominant component by binauralizing at a spatial location relative to an intended listener based on the dominant audio component direction to form a rendered binauralized estimated dominant component; (c) reconstructing a residual component estimate from the first (e.g., initial) output representation; and (d) combining the rendered binauralized estimated dominant component and the residual component estimate to form an output spatialized audio encoded signal.
[0028] The encoded audio signal may further comprise a series of residual matrix coefficients representing the residual audio signal, and step (c) may further comprise: (c1) applying the residual matrix coefficients to the first (eg, initial) output representation to reconstruct the residual component estimate.
[0029] In some embodiments, the residual component estimate may be reconstructed by subtracting the rendered binauralized estimated dominant component from the first (e.g., initial) output representation.Step (b) may comprise performing an initial rotation of the estimated dominant component in accordance with an input head tracking signal indicative of the head orientation of the intended listener.
[0030] According to a further aspect of the invention, there is provided a method for decoding and reproducing an audio stream for a listener using headphones, the method comprising: (a) receiving a data stream comprising a first audio representation and additional audio transformation data; (b) receiving head position data representing the position of the listener; (c) creating one or more auxiliary signals based on the first audio representation and the received transformation data; (d) creating a second audio representation, the second audio representation comprising a combination of the first audio representation and the auxiliary signal(s), in which one or more of the auxiliary signal(s) have been modified in response to the head position data; and (e) outputting the second audio representation as an output audio stream.
[0031] In some embodiments, the modification of the auxiliary signal may further include a simulation of the acoustic path from the sound source position to the listener's ear. The transform data may include matrixed coefficients and at least one of the following: the sound source position or the sound source direction. The transform processing may be applied according to time or frequency. The auxiliary signal may represent at least one dominant component. The sound source position or direction may be received as part of the transform data and may be rotated in response to the head orientation data. In some embodiments, the maximum rotation amount is limited to a value less than 360 degrees in azimuth or elevation. The second representation may be obtained from the first representation by matrixing in a transform domain or a filter bank domain. The transform data may further include additional matrixed coefficients, and step (d) may further include modifying the first audio representation in response to the additional matrixed coefficients before combining the first audio representation and (one or more) auxiliary audio signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Embodiments of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0033] Figure 1 schematically illustrates a headphone decoder for matrix-encoded content;
[0034] Figure 2 schematically illustrates an encoder according to an embodiment;
[0035] Figure 3 is a schematic block diagram of a decoder;
[0036] Figure 4 is a detailed visualization of the encoder; and
[0037] Figure 5 One form of decoder is shown in more detail. DETAILED DESCRIPTION
[0038] Embodiments provide a system and method for representing object-based or channel-based audio content that (1) is compatible with stereo playback, (2) enables binaural playback including head tracking, (3) has low decoder complexity, and (4) does not rely on matrix coding, but is still compatible with matrix coding.
[0039] This is achieved by combining an encoder-side analysis of one or more dominant components (or dominant objects or a combination thereof) including weights for predicting these dominant components from the downmix in combination with additional parameters, the weights minimizing the error between a binaural rendering based solely on the leading or dominant components and a desired binaural representation of the entire content.
[0040] In an embodiment, the analysis of the dominant component (or components) is provided in the encoder rather than in the decoder / renderer. The audio stream is then supplemented with metadata indicating the direction of the dominant component and information about how the dominant component(s) can be obtained from the associated downmix signal.
[0041] Figure 2 A form of encoder 20 of a preferred embodiment is shown. Object-based or channel-based content 21 is analyzed 23 to determine (one or more) dominant components. The analysis can occur in terms of time and frequency (assuming that the audio content is decomposed into time slices and frequency sub-slices). The result of this processing is a dominant component signal 26 (or multiple dominant component signals) and associated (one or more) position or (one or more) direction information 25. Subsequently, weights are estimated 24 and output 27 so that (one or more) dominant component signals can be reconstructed from the transmitted downmix. The downmix generator 22 does not necessarily have to comply with the LtRt downmix rule, but can be a standard ITU (LoRo) downmix using non-negative real-valued downmix coefficients. Finally, the output downmix signal 29, weights 27 and position data 25 are packaged by an audio encoder 28 and are ready for distribution.
[0042] Now go to Figure 3 , showing the corresponding decoder 30 of the preferred embodiment. The audio decoder reconstructs the downmix signal. This signal is input 31 and unpacked by the audio decoder 32 into the downmix signal, the directions and weights of the dominant components. Subsequently, the dominant component estimated weights are used to reconstruct (34) the (one or more) leading components, which are rendered 36 using the sent position or direction data. The position data can be optionally modified 33 based on head rotation or translation information 38. In addition, the (one or more) reconstructed dominant components can be subtracted (35) from the downmix. Optionally, the subtraction of the (one or more) dominant components occurs within the downmix path, but alternatively, as described below, this subtraction can also take place at the encoder.
[0043] To improve the removal or cancellation of the reconstructed dominant component in the subtractor 35, the dominant component output may first be rendered using the transmitted position or orientation data before the subtraction. Figure 3 This optional rendering stage 39 is shown in .
[0044] Now let's go back to the beginning and describe the encoder in more detail. Figure 4One form of an encoder 40 for processing object-based (e.g., Dolby Atmos) audio content is shown. The audio objects are initially stored as Atmos objects 41 and are first divided into time slices and frequency slices using a hybrid complex-valued quadrature mirror filter (HCQMF) bank 42. When we omit the corresponding time index and frequency index, the input object signal can be represented by x i [n] represents the corresponding position in the current frame; the corresponding position in the current frame is represented by the unit vector Given, index i refers to the object number and index n refers to the time (e.g., subband sample index). Input object signal x i [n] is an example of channel-based or object-based input audio.
[0045] Using the complex-valued scalar H l,i 、H r,i (e.g., single-tap HRTF 48) to create 43 anechoic sub-band binaural mix Y(y l ,y r ), complex-valued scalar H l,i 、H r,i Indicates the position corresponding to The sub-band representation of HRIR is:
[0046]
[0047]
[0048] Alternatively, a binaural mix Y(y) can be created by using the head-related impulse response (HRIR) l ,y r ). In addition, the amplitude shift gain coefficient g is used l,i 、g r,i To create a 44 stereo downmix l 、z r (Exemplary implementation of initial output representation):
[0049]
[0050]
[0051] The dominant component can be estimated by The dominant component 45 is calculated by first computing a weighted sum of the unit direction vectors for each object:
[0052]
[0053] in, is the signal x iEnergy of [n]:
[0054]
[0055] in,(.) * is the complex conjugate operator.
[0056] The dominant / pilot signal d[n] (exemplarily realizing the dominant audio component) is then given by the following equation:
[0057]
[0058] in, is generated with the unit vector For example, to create a virtual microphone with a directional pattern based on higher-order spherical harmonics, one implementation would correspond to:
[0059]
[0060] in, represents a unit direction vector in a two-dimensional or three-dimensional coordinate system, (.) represents the dot product operator of two vectors, and a, b, and c represent exemplary parameters (eg, a=b=0.5; c=1).
[0061] Calculate 46 weights or prediction coefficients w l,d 、w r,d , and using these weights or prediction coefficients w l,d 、w r,d To calculate the estimated guidance signal 47
[0062]
[0063] Among them, the weight w l,d 、w r,d Minimize the downmix signal z l 、z r Given d[n] and The mean square error between the weights w l,d 、w r,d is used to represent the initial output (e.g., z l 、z r ) is mapped to the dominant audio component (e.g. ). A known method of deriving these weights is by applying a minimum mean square error (MMSE) predictor:
[0064]
[0065] Among them, R abis the covariance matrix between the signals for signal a and signal b, and ∈ is the regularization parameter.
[0066] We can then extract the anechoic binaural mix y l 、y r Subtract 49 dominant component signal The rendering of the estimate in order to use the dominant component signal Direction / position Associated HRTF (HRIR) H l,D 、H r,D 50 to create residual binaural mix
[0067]
[0068]
[0069] Finally, another set of prediction coefficients or weights w is estimated 51 i,j , these prediction coefficients or weights w i,j This allows for the minimum mean square error estimation to be used to estimate the stereo mix z l 、z r Reconstructing the residual binaural mix
[0070]
[0071] Among them, R ab is the covariance matrix between the signals representing a and b, and ∈ is the regularization parameter. The prediction coefficients or weights w i,j is used to represent the initial output (e.g., z l 、z r ) is mapped to the residual binaural mix An example of the estimated residual matrix coefficients of . Additional level constraints can be imposed on the above expression to overcome any prediction loss. The encoder outputs the following information:
[0072] Stereo Mix l 、z r (exemplarily implementing the initial output representation);
[0073] Estimate the coefficient w of the dominant component l,d 、w r,d (exemplarily implementing the dominant audio component weighting factor);
[0074] The position or direction of the dominant component
[0075] and optionally, the residual weight w i,j (Exemplarily implementing the residual matrix coefficients).
[0076] Although the above description relates to rendering based on a single dominant component, in some embodiments, the encoder may be adapted to detect multiple dominant components, determine a weight and direction for each of the multiple dominant components, render each of the multiple dominant components and subtract each of the multiple dominant components from the anechoic binaural mix Y, and then determine the residual weight after each of the multiple dominant components has been subtracted from the anechoic binaural mix Y.
[0077] Decoder / Renderer
[0078] Figure 5 One form of decoder / renderer 60 is shown in more detail. The decoder / renderer 60 applies a decoder / renderer 60 which is intended to be used with input information z which has not been unpacked. l 、z r ;w l,d 、w r,d ; w i,j Reconstructing binaural mix l 、y r For processing to be output to the listener 71. Here, the stereo mix z l 、z r is an example of a first audio representation, and the prediction coefficients or weights w i,j and / or dominant component signal Direction / position This is an example of additional audio conversion data.
[0079] First, the stereo downmix is divided into time / frequency slices using a suitable filter bank or transform 61, such as the HCQMF analysis bank 61. Other transforms such as the discrete Fourier transform, (modified) cosine or sine transform, time domain filter bank or wavelet transform can be equally applied. Subsequently, the prediction coefficient weights w are used. l,d 、w r,d To calculate the estimated dominant component signal
[0080]
[0081] Estimated dominant component signal is an example of an auxiliary signal. Hence, this step may be said to correspond to creating one or more auxiliary signals based on said first audio representation and the received transformation data.
[0082] The dominant component signal is then rendered 65 and based on the position / orientation data transmitted HRTF 69 is modified 68, the position / orientation data sent It may be modified (rotated) based on information obtained from the head tracker 62. Finally, the total mute binaural output contains the weights w based on the prediction coefficients. i,j The reconstructed residual The dominant component signal of summation 66:
[0083]
[0084]
[0085] The total anonymized binaural output is an example of a second audio representation. Thus, this step may be said to correspond to creating a second audio representation comprising a combination of the first audio representation and the auxiliary signal(s), wherein one or more of the auxiliary signal(s) have been modified in response to the head position data.
[0086] It should further be noted that if information about more than one dominant signal is received, each dominant signal may be rendered and added to the reconstructed residual signal.
[0087] As long as head rotation or translation is not applied, the output signal should be very close (in terms of RMS error) to the reference binaural signal y l 、y r ,if only
[0088]
[0089] Key properties
[0090] As can be observed from the above equation formulation, the effective operation to construct the anechoic binaural representation from the stereo representation involves a 2x2 matrix 70 where the matrix coefficients depend on the transmitted information w l,d 、w r,d ; w i,j and head tracker rotations and / or translations. This indicates that the processing complexity is relatively low, since the analysis of the dominant component is applied in the encoder rather than in the decoder.
[0091] If the dominant component is not estimated (e.g., w l,d 、w r,d =0), the described solution is equivalent to the parameterized binaural approach.
[0092] In the case where it is desired to exclude certain objects from head rotation / head tracking, these objects can be excluded from (1) the dominant component direction analysis and (2) the dominant component signal prediction. As a result, these objects will be excluded by the coefficient w i,jConverts from stereo to binaural so it is not affected by any head rotation or translation.
[0093] In a similar vein, objects can be set to "through" mode, which means that they will be amplitude-shifted in the binaural representation instead of HRIR convolved. This can be done by simply changing the coefficients This can be achieved using amplitude panning gains instead of single-tap HRTFs, or using any other suitable binaural processing.
[0094] Extensions
[0095] Embodiments are not limited to the use of stereo downmix, as other channel counts may also be employed.
[0096] Reference Figure 5 The decoder 60 described has an output signal containing the rendered dominant component direction plus the matrix coefficients w i,j The matrixed input signal. The coefficients can be derived in various ways, for example:
[0097] 1. Can be used in the encoder with the help of signal The parameterized reconstruction is used to determine the coefficient w i,j In other words, in this implementation, the coefficient w i,j Aims to faithfully reconstruct the binaural signal y l 、y r , these binaural signals would be what would have been obtained when binaurally rendering the original input objects / channels; in other words, the coefficients w i,j It is content-driven.
[0098] 2. The coefficient w can be i,j The coefficients representing the HRTF are sent from the encoder to the decoder to represent the HRTF for a fixed spatial position (e.g., a spatial position at + / - 45 degrees in azimuth). In other words, the residual signal is processed to simulate the reproduction on two virtual loudspeakers at certain locations. When these coefficients representing the HRTF are sent from the encoder to the decoder, the locations of the virtual loudspeakers may change with time and frequency. If this method is used to represent the residual signal by using static virtual loudspeakers, the coefficients w i,j does not need to be sent from the encoder to the decoder and can instead be hardwired in the decoder. A variation of this approach would involve a limited set of static positions available in the decoder with their corresponding coefficients w i,j , and the choice of which static position to use for processing the residual signal is signaled from the encoder to the decoder.
[0099] Signal More than two signals can be reconstructed by means of a statistical analysis of these signals at the decoder followed by binaural rendering of the resulting upmixed signals, via a so-called upmixer.
[0100] The described method can also be applied to systems where the transmitted signal Z is a binaural signal. In this specific case, Figure 5 The decoder 60 remains as is, while Figure 4 The block 44 labeled "Generate Stereo (LoRo) Mix" in the example should be replaced by the same "Generate Anechoic Binaural Mix" 43 ( Figure 4 ) are replaced. In addition, other forms of mixes can be generated upon request.
[0101] The method can be extended to methods for reconstructing one or more FDN input signals from a transmitted stereo mix containing specific objects or channel subsets.
[0102] This method can be extended to predict multiple dominant components from the transmitted stereo mix and render these dominant components at the decoder. There is no restriction on predicting only one dominant component per time / frequency slice. Specifically, the number of dominant components can be different in each time / frequency slice.
[0103] explain
[0104] Reference throughout this specification to "one embodiment," "some embodiments," or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Furthermore, in one or more embodiments, the particular features, structures, or characteristics may be combined in any suitable manner as would be apparent to one of ordinary skill in the art from this disclosure.
[0105] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc., to describe common objects indicates only that different instances of similar objects are being referred to and is not intended to imply that the objects so described must be in a given order in time, space, ranking, or in any other manner.
[0106] In the appended claims and the description herein, any of the terms "comprise", "it includes" or "it includes" is an open term that means including at least the following elements / features, but does not exclude other elements / features. Therefore, the term "comprise" when used in a claim should not be interpreted as limiting the means or elements or steps listed thereafter. For example, the scope of the expression "a device including A and B" should not be limited to a device consisting of only elements A and B. As used herein, any of the terms "comprise" or "it includes" or "it includes" is also an open term that means including at least the elements / features following the term, but does not exclude other elements / features. Therefore, include is synonymous with include and means include.
[0107] As used herein, the term "exemplary" is used in the sense of providing an example, as opposed to an indicative quality. That is, an "exemplary embodiment" is an embodiment provided as an example, as opposed to an embodiment that is necessarily of exemplary quality.
[0108] It should be appreciated that in the above description of exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than those expressly recited in each claim. On the contrary, as reflected in the appended claims, inventive aspects lie in fewer features than all the features of a single aforementioned disclosed embodiment. Therefore, the claims appended to the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0109] Furthermore, while some embodiments described herein include some, but not others, features included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the present invention and form different embodiments as will be understood by those skilled in the art. For example, in the appended claims, any of the claimed embodiments may be used in any combination.
[0110] In addition, some of the embodiments are described herein as methods or combinations of elements that can be implemented by a processor of a computer system or by other means for implementing a function. Therefore, a processor with instructions required for implementing such a method or an element of a method forms a means for implementing the method or an element of the method. In addition, the elements described herein of the device embodiments are examples of means for implementing the functions performed by the elements for implementing the purpose of the present invention.
[0111] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the present invention may be practiced without these specific details. In other cases, well-known methods, structures, and techniques are not presented in detail in order not to obscure the understanding of the description.
[0112] Similarly, it should be noted that the term "coupled" when used in the claims should not be interpreted as being limited to only direct connections. The terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms are not intended to be synonyms of each other. Thus, the scope of the expression "device A coupled to device B" should not be limited to devices or systems in which the output of device A is directly connected to the input of device B. It means that there is a path between the output of A and the input of B, which path may be a path that includes other devices or means. "Coupled" can mean that two or more elements are in direct physical or electrical contact, or that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0113] Thus, while embodiments of the present invention have been described, those skilled in the art will recognize that other and further modifications may be made to the present invention without departing from the spirit of the present invention, and it is intended that all such changes and modifications falling within the scope of the present invention be claimed. For example, any formulas given above are merely representative of processes that may be used. Functionality may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to the methods described within the scope of the present invention.
[0114] Aspects of the present invention may be appreciated from the following enumerated example embodiments (EEES):
[0115] EEE 1. A method for encoding channel-based or object-based input audio for playback, the method comprising the steps of:
[0116] (a) First, the channel-based or object-based input audio is rendered into an initial output representation;
[0117] (b) determining an estimate of a dominant audio component from the channel-based or object-based input audio, and determining a series of dominant audio component weighting factors for mapping the initial output representation to the dominant audio component;
[0118] (c) determining an estimate of the direction or position of the dominant audio component; and
[0119] (d) Encoding the initial output representation, the dominant audio component weighting factor, the dominant audio component direction or position into an encoded signal for playback.
[0120] EEE 2. The method of EEE 1, further comprising determining an estimate of a residual mix, the residual mix being a rendering of the initial output representation minus a dominant audio component or an estimate of the dominant audio component. EEE 2. The method of claim 1, further comprising determining an estimate of a residual mix, the residual mix being a rendering of the initial output representation minus a dominant audio component or an estimate of the dominant audio component.
[0121] EEE 3. The method according to EEE 1, further comprising: generating an anechoic binaural mix of the channel-based or object-based input audio, and determining an estimate of the residual mix, wherein the estimate of the residual mix is a rendering of the anechoic binaural mix minus a dominant audio component or an estimate of the dominant audio component.
[0122] EEE 4. The method according to EEE 2 or 3, further comprising determining a series of residual matrix coefficients for mapping the initial output representation to an estimate of the residual mix. EEE 4.
[0123] EEE 5. The method according to any preceding EEE, wherein the initial output representation comprises a headphone or loudspeaker representation. ...
[0124] EEE 6. The method according to any preceding EEE, wherein the channel-based or object-based input audio is sliced in time and frequency, and the encoding step is repeated for a series of time steps and a series of frequency bands. EEE 6. The method according to any preceding EEE, wherein the channel-based or object-based input audio is sliced in time and frequency.
[0125] EEE 7. The method according to any preceding EEE, wherein the initial output representation comprises a stereo speaker mix.
[0126] EEE 8. A method for decoding an encoded audio signal, the encoded audio signal comprising:
[0127] - first output representation;
[0128] - dominant audio component direction and dominant audio component weighting factor;
[0129] The method comprises the following steps:
[0130] (a) determining an estimated dominant component using the dominant audio component weighting factor and the initial output representation;
[0131] (b) rendering the estimated dominant component by binauralizing at a spatial location relative to an intended listener according to the dominant audio component directions to form a rendered binauralized estimated dominant component;
[0132] (c) reconstructing the residual component estimate from the first output representation; and
[0133] (d) Combining the rendered binauralized estimated dominant component and the residual component estimate to form an output spatialized audio coded signal.
[0134] EEE 9. The method according to EEE 8, wherein the encoded audio signal further comprises a series of residual matrix coefficients representing the residual audio signal, and step (c) further comprises:
[0135] (c1) Applying the residual matrix coefficients to the first output representation to reconstruct the residual component estimate.
[0136] EEE 10. The method of EEE 8, wherein the residual component estimate is reconstructed by subtracting a dominant component of the rendered binauralized estimate from the first output representation. ...
[0137] EEE 11. The method of EEE 8, wherein step (b) comprises performing an initial rotation of the estimated dominant component based on an input head tracking signal indicative of a head orientation of the intended listener. EEE 12. The method of claim 11, wherein step (b) comprises performing an initial rotation of the estimated dominant component based on an input head tracking signal indicative of a head orientation of the intended listener.
[0138] EEE 12. A method for decoding and reproducing an audio stream for a listener using headphones, the method comprising:
[0139] (a) receiving a data stream comprising a first audio representation and additional audio transform data;
[0140] (b) receiving head position data indicating the position of the listener;
[0141] (c) creating one or more auxiliary signals based on the first audio representation and the received transformation data;
[0142] (d) creating a second audio representation, the second audio representation comprising a combination of the first audio representation and the auxiliary signal(s), in which one or more of the auxiliary signal(s) has been modified in response to the head position data; and
[0143] (e) Outputting the second audio representation as an output audio stream.
[0144] EEE 13. The method according to EEE 12, wherein the modification of the auxiliary signal comprises a simulation of an acoustic path from the sound source position to the listener's ear.
[0145] EEE 14. The method according to EEE 12 or 13, wherein the transformation data comprises matrixed coefficients and at least one of: a sound source position or a sound source direction. EEE 14. The method according to EEE 12 or 13, wherein the transformation data comprises matrixed coefficients and at least one of the following: a sound source position or a sound source direction.
[0146] EEE 15. The method according to any one of EEEs 12 to 14, wherein the transform process is applied according to time or frequency. EEE 15.
[0147] EEE 16. The method according to any one of EEEs 12 to 15, wherein the auxiliary signal represents at least one dominant component. EEE 16.
[0148] EEE 17. A method according to any one of EEEs 12 to 16, wherein a sound source position or direction received as part of the transformation data is rotated in response to the head orientation data. EEE 17. A method according to any one of EEEs 12 to 16, wherein the sound source position or direction received as part of the transformation data is rotated in response to the head orientation data.
[0149] EEE 18. The method of EEE 17, wherein the maximum rotation amount is limited to a value less than 360 degrees in azimuth or elevation.
[0150] EEE 19. The method according to any one of EEEs 12 to 18, wherein the second representation is obtained from the first representation by matrixing in a transform domain or a filter bank domain. ...
[0151] EEE 20. A method according to any one of EEEs 12 to 19, wherein the transform data further comprises additional matrixed coefficients, and step (d) further comprises modifying the first audio representation in response to the additional matrixed coefficients before combining the first audio representation and the auxiliary audio signal(s).
[0152] EEE 21. An apparatus comprising one or more devices configured to perform the method according to any one of EEEs 1 to 20. EEE 22. An apparatus comprising:
[0153] EEE 22. A computer-readable storage medium comprising an instruction program that, when executed by one or more processors, causes one or more devices to perform the method according to any one of EEEs 1 to 20. EEE 23.
Claims
1. A system configured to encode channel-based or object-based input audio (21) for playback, comprising: one or more processors; as well as A computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: Rendering the channel-based or object-based input audio (21) into an initial output representation; An estimate of a dominant audio component (26) is determined (23) from channel-based or object-based input audio (21), the determination comprising: determining (24) a series of dominant audio component weighting factors (27) for mapping the initial output representation into dominant audio components, and determining an estimate of a dominant audio component (26) based on the dominant audio component weighting factor (27) and the initial output representation; determining an estimate of the direction or position (25) of the dominant audio component; and At least one of a dominant audio component direction or position, an initial output representation, and a dominant audio component weighting factor (27) are encoded into an encoded signal for playback. 2 . The system of claim 1 , the operations further comprising determining an estimate of a residual mix, the residual mix being a rendering of the initial output representation minus a dominant audio component or an estimate of the dominant audio component.
3. The system of claim 1 , the operations further comprising generating an anechoic binaural mix of the input audio, which is channel-based or object-based, and determining an estimate of the residual mix, wherein The estimate of the residual mix is a rendering of the anechoic binaural mix minus the dominant audio component or an estimate of the dominant audio component. 4 . The system of claim 2 , the operations further comprising determining a series of residual matrix coefficients for mapping the initial output representation to an estimate of the residual mix.
5. The system according to any one of claims 1 to 4, wherein: The initial output representation includes a headphone representation or a loudspeaker representation.
6. The system according to any one of claims 1 to 4, wherein: The channel-based or object-based input audio is sliced in time and frequency, and the encoding step is repeated for a series of time steps and a series of frequency bands.
7. The system according to any one of claims 1 to 4, wherein: The initial output representation comprises a stereo speaker mix.
Citation Information
Cited By
Head tracking for parameterized binaural output systems and methods
CN121151789A