Convert binaural signals into stereo audio signals
By analyzing the band direction parameters of binaural audio signals, modifying the differences between channels and applying spectrum adjustments, a stereo audio signal suitable for speaker reproduction is solved, and the spatial perception and tone perception differences when binaural signals are converted into stereo audio signals is achieved, achieving high-quality audio reproduction.
Patent Information
- Application Number
- CN202080081512.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-25
- Filing Date
- 2020-11-13
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-11-13
AI Technical Summary
The prior art is difficult to effectively convert binaural signals into stereo audio signals, resulting in differences in spatial perception and tone perception when speakers are reproduced.
By analyzing the band direction parameters of binaural audio signals, modifying the differences between channels, generating left and right channel audio signals, and applying spectrum adjustment and head-related transfer functions to compensate, a stereo audio signal suitable for speaker reproduction is generated.
Accurate spatial perception and colorless tone perception during speaker reproduction are achieved, reducing spatial and tone perception differences of binaural signals and improving audio quality.
Smart Images

Figure CN114762040B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to apparatuses and methods for converting binaural signals into stereophonic audio signals, but non-exclusively to apparatuses and methods for performing the conversion within a spatial audio signal environment. Background Art
[0002] Human perception of sound direction is based on binaural cues, which include the interaural time difference (ITD), the interaural level difference (ILD), and spectral cues. Amplitude panning (e.g., VBAP as discussed by Ville Pulkki in "Virtual Sound Source Positioning Using Vector Base Amplitude Panning" (Journal of the Audio Engineering Society, 1997)) is commonly used to generate stereophonic signals for loudspeaker reproduction, and when the amplitude-panned sound is reproduced by stereophonic loudspeakers and listened to by a human listener, the amplitude panning transforms into these cues.
[0003] Accordingly, human perception of the spaciousness and envelopment of sound is based on binaural cues related to the interaural coherence (IC). Stereophonic signals are typically generated in such a way (e.g., using a reverberator) that when the stereophonic signal is reproduced by stereophonic loudspeakers, IC cues that generate a perception such as width or spaciousness are generated at the human ears.
[0004] On the other hand, binaural signals are meant to be reproduced by headphones. Therefore, binaural cues (including ITD, ILD, IC, and spectral cues) need to be inherent in the audio signal itself. This can be achieved, for example, by recording spatial sound with a microphone at the entrance of the ear canal of a real human or an artificial head. It can also be achieved, for example, by synthetically generating binaural sound by applying appropriate head-related transfer functions (HRTFs) and a reverberator to a multi-channel loudspeaker mix. When such a binaural recording (or generally, binaural audio) is reproduced by headphones (possibly after headphone calibration), a true perception of spatial sound is achieved.
[0005] Immersive audio codecs are being implemented to support a wide range of operating points from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Service (IVAS) codec, which is designed to be suitable for use on communication networks such as 3GPP 4G / 5G networks, including use in immersive services such as immersive voice and audio for virtual reality (VR). <>
[0006] The input signal can be presented to the IVAS encoder in one of a variety of supported formats (and in some allowed format combinations).
[0007] It has been proposed that IVAS uses binaural signals as input and has a conventional stereo audio output.
[0008] Since stereo audio signals are more suitable for speaker playback, there is a need for devices and methods for effectively converting binaural signals into conventional stereo audio signals. SUMMARY OF THE INVENTION
[0009] According to a first aspect, there is provided an apparatus comprising components configured to perform the following operations: obtaining a binaural audio signal; based on the binaural audio signal, obtaining at least one direction parameter of at least one frequency band of the binaural audio signal; based on the at least one direction parameter of the at least one frequency band, processing the binaural audio signal by modifying the inter-channel difference of at least one frequency band of the binaural audio signal to generate at least two audio signals for speaker reproduction; and outputting the at least two audio signals for speaker reproduction.
[0010] The inter-channel difference of at least one frequency band of the binaural audio signal may include at least one of the following: at least one energy / amplitude difference of the channels of the binaural audio signal; at least one phase difference of the channels of the binaural audio signal; and at least one time difference of the channels of the binaural audio signal.
[0011] The component configured to process the binaural audio signal to generate at least two audio signals for speaker reproduction may be configured to: further apply spectral adjustment to the processed at least one frequency band based on the at least one direction parameter of the at least one frequency band.
[0012] The component configured to process the binaural audio signal by modifying the inter-channel difference of at least one frequency band based on the at least one direction parameter of the at least one frequency band to generate at least two audio signals for speaker reproduction may be configured to: generate an estimate of at least a part of a covariance matrix for at least one frequency band of the binaural audio signal; generate an energy estimate for at least one frequency band of the binaural audio signal; generate at least a part of a target covariance matrix for at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band; generate a mixing matrix for mixing at least one frequency band of the binaural audio signal; and generate a left-channel audio signal and a right-channel audio signal from the channel combination of at least one frequency band of the binaural audio signal based on the mixing matrix.
[0013] The at least two audio signals for speaker reproduction may include a left-channel audio signal and a right-channel audio signal.
[0014] A component configured to process a binaural audio signal to generate at least two audio signals for loudspeaker reproduction by modifying the inter-channel differences of at least one frequency band based on at least one direction parameter of at least one frequency band may be further configured to: for at least one frequency band, generate a decorrelated audio signal based on the binaural audio signal; for the decorrelated audio signal, generate another mixing matrix; based on the another mixing matrix, generate another left-channel audio signal and another right-channel audio signal from channel combinations of at least one frequency band of the decorrelated audio signal; combine the left-channel audio signal and the another left-channel audio signal to generate a combined left channel; and combine the right-channel audio signal and the another right-channel audio signal to generate a combined right channel, and wherein the at least two audio signals for loudspeaker reproduction include the combined left-channel audio signal and the combined right-channel audio signal.
[0015] A component configured to further apply spectral adjustment to the processed at least one frequency band based on at least one direction parameter of at least one frequency band may be configured to: determine a binaural response and / or a long-term response estimate based on the direction parameter of at least one frequency band; and compensate for the determined binaural response and / or long-term response estimate from the processed at least one frequency band.
[0016] The binaural response and / or the long-term response may include at least one of the following: at least one energy / amplitude; at least one correlation of channels of the binaural audio signal; at least one phase difference of channels of the binaural audio signal; at least one time difference of channels of the binaural audio signal.
[0017] The binaural response and / or the long-term response may include the spectrum of the binaural audio signal, and wherein a component configured to remove the determined binaural response and / or long-term response estimate from the processed at least one frequency band may be configured to: obtain a filter and / or a gain based on the estimated direction parameter and an average head-related transfer function corresponding to the at least one direction parameter; and apply the filter and / or the gain to the processed at least one frequency band.
[0018] A component configured to determine a binaural response and / or a long-term response estimate based on the direction parameter of at least one frequency band may be configured to: generate a long-term equalization filter by comparing an average spectrum of the binaural signals with a predetermined HRTF data set, and wherein a component configured to remove the determined binaural response and / or long-term response estimate from the processed at least one frequency band may be configured to: apply the long-term equalization filter to the processed at least one frequency band.
[0019] A component configured to obtain at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal may be configured to: analyze at least one frequency band of the binaural audio signal to determine at least one direction parameter of at least one frequency band.
[0020] A component configured to analyze at least one frequency band of a binaural audio signal to determine at least one direction parameter of the at least one frequency band may be further configured to: for the at least one frequency band, estimate a delay that maximizes the correlation between channels of the binaural audio signal; and formulate a direction parameter based on the estimated delay.
[0021] The above component may be further configured to: for at least one frequency band of the binaural audio signal, obtain a direct-to-total energy ratio value based on the measured normalized correlation between channels of the binaural audio signal.
[0022] A component configured to generate at least a portion of a target covariance matrix for at least one frequency band of a binaural audio signal based on at least one direction parameter of the at least one frequency band may be further configured to: for at least one frequency band of the binaural audio signal, further generate at least a portion of the target covariance matrix based on the direct-to-total energy ratio value of the at least one frequency band.
[0023] A component configured to determine a binaural response and / or a long-term response estimate based on at least one direction parameter of at least one frequency band may be configured to: determine the binaural response and / or the long-term response estimate based on the direct-to-total energy ratio value of the at least one frequency band.
[0024] A component configured to obtain a binaural audio signal may be configured to perform one of the following: capture a binaural audio signal with an artificial head; capture a binaural audio signal at the entrance of a user's ear canal; render a binaural audio signal according to a head-related transfer function; and render a binaural audio signal using a binaural room impulse response.
[0025] A component configured to output at least two audio signals for speaker reproduction may be configured to: output the at least two audio signals for speaker reproduction to a stereo speaker.
[0026] According to a second aspect, there is provided a method, which includes: obtaining a binaural audio signal; based on the binaural audio signal, obtaining at least one direction parameter of at least one frequency band of the binaural audio signal; based on the at least one direction parameter of the at least one frequency band, processing the binaural audio signal by modifying the inter-channel difference of at least one frequency band of the binaural audio signal to generate at least two audio signals for speaker reproduction; and outputting the at least two audio signals for speaker reproduction.
[0027] The inter-channel differences of at least one frequency band of a binaural audio signal may include at least one of the following: at least one energy / amplitude difference of channels of the binaural audio signal; at least one phase difference of channels of the binaural audio signal; and at least one time difference of channels of the binaural audio signal.
[0028] Processing the binaural audio signal to generate at least two audio signals for speaker reproduction may include: further applying spectral adjustment to the processed at least one frequency band, further based on at least one direction parameter of at least one frequency band.
[0029] Processing the binaural audio signal to generate at least two audio signals for speaker reproduction by modifying the inter-channel differences of at least one frequency band based on at least one direction parameter of at least one frequency band may include: generating an estimate of at least a part of a covariance matrix for at least one frequency band of the binaural audio signal; generating an energy estimate for at least one frequency band of the binaural audio signal; generating at least a part of a target covariance matrix for at least one frequency band of the binaural audio signal based on at least one direction parameter of at least one frequency band; generating a mixing matrix for at least one frequency band for mixing the binaural audio signal; and generating a left-channel audio signal and a right-channel audio signal from a combination of channels of at least one frequency band of the binaural audio signal based on the mixing matrix.
[0030] The at least two audio signals for speaker reproduction may include a left-channel audio signal and a right-channel audio signal.
[0031] Processing the binaural audio signal to generate at least two audio signals for speaker reproduction by modifying the inter-channel differences of at least one frequency band based on at least one direction parameter of at least one frequency band may include: generating a decorrelated audio signal for at least one frequency band based on the binaural audio signal; generating another mixing matrix for the decorrelated audio signal; generating another left-channel audio signal and another right-channel audio signal from a combination of channels of at least one frequency band of the decorrelated audio signal based on the another mixing matrix; combining the left-channel audio signal and the another left-channel audio signal to generate a combined left channel; and combining the right-channel audio signal and the another right-channel audio signal to generate a combined right channel, and wherein the at least two audio signals for speaker reproduction include the combined left-channel audio signal and the combined right-channel audio signal.
[0032] Further applying spectral adjustment to the processed at least one frequency band, further based on at least one direction parameter of at least one frequency band, may include: determining a binaural response and / or a long-term response estimate based on the direction parameter of at least one frequency band; and compensating for the determined binaural response and / or long-term response estimate from the processed at least one frequency band.
[0033] The binaural response and / or the long-term response may include at least one of the following: at least one energy / amplitude; at least one correlation of the channels of the binaural audio signal; at least one phase difference of the channels of the binaural audio signal; at least one time difference of the channels of the binaural audio signal.
[0034] The binaural response and / or the long-term response may include the spectrum of the binaural audio signal, and wherein removing the determined binaural response and / or long-term response estimate from at least one processed frequency band may include: obtaining a filter and / or gain based on the estimated direction parameter and the average head-related transfer function corresponding to at least one direction parameter; and applying the filter and / or gain to at least one processed frequency band.
[0035] Determining the binaural response and / or long-term response estimate based on the direction parameter of at least one frequency band may include: generating a long-term equalization filter by comparing the average spectrum of the binaural signals with a predetermined HRTF data set, and wherein removing the determined binaural response and / or long-term response estimate from at least one processed frequency band may include: applying the long-term equalization filter to at least one processed frequency band.
[0036] Obtaining at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal may include: analyzing at least one frequency band of the binaural audio signal to determine at least one direction parameter of at least one frequency band.
[0037] Analyzing at least one frequency band of the binaural audio signal to determine at least one direction parameter of at least one frequency band may include: for at least one frequency band, estimating a delay that maximizes the correlation between the channels of the binaural audio signal; and formulating a direction parameter based on the estimated delay.
[0038] The method may further include: for at least one frequency band of the binaural audio signal, obtaining a direct-to-total energy ratio based on the measured normalized correlation between the channels of the binaural audio signal.
[0039] Generating at least a portion of the target covariance matrix for at least one frequency band of the binaural audio signal based on at least one direction parameter of at least one frequency band may further include: for at least one frequency band of the binaural audio signal, further generating at least a portion of the target covariance matrix based on the direct-to-total energy ratio of at least one frequency band.
[0040] Determining the binaural response and / or long-term response estimate based on at least one direction parameter of at least one frequency band may include: determining the binaural response and / or long-term response estimate based on the direct-to-total energy ratio of at least one frequency band.
[0041] Obtaining a binaural audio signal may include performing one of the following: capturing a binaural audio signal with an artificial head; capturing a binaural audio signal at the entrance of a user's ear canal; rendering a binaural audio signal according to a head-related transfer function; and rendering a binaural audio signal using a binaural room impulse response.
[0042] Outputting at least two audio signals for speaker reproduction may include: outputting at least two audio signals for speaker reproduction to stereo speakers.
[0043] According to a third aspect, there is provided an apparatus, which includes at least one processor and at least one memory including computer program code, and the at least one memory and the computer program code are configured to, together with the at least one processor, cause the apparatus to at least: obtain a binaural audio signal; based on the binaural audio signal, obtain at least one direction parameter of at least one frequency band of the binaural audio signal; based on at least one direction parameter of at least one frequency band, process the binaural audio signal by modifying an inter-channel difference of at least one frequency band of the binaural audio signal to generate at least two audio signals for speaker reproduction; and output at least two audio signals for speaker reproduction.
[0044] The inter-channel difference of at least one frequency band of the binaural audio signal may include at least one of the following: at least one energy / amplitude difference between channels of the binaural audio signal; at least one phase difference between channels of the binaural audio signal; and at least one time difference between channels of the binaural audio signal.
[0045] The apparatus that is caused to process the binaural audio signal to generate at least two audio signals for speaker reproduction may be caused to: further apply spectral adjustment to the processed at least one frequency band further based on at least one direction parameter of at least one frequency band.
[0046] The apparatus that is caused to process the binaural audio signal by modifying an inter-channel difference of at least one frequency band based on at least one direction parameter of at least one frequency band to generate at least two audio signals for speaker reproduction may be caused to: generate an estimate of at least a part of a covariance matrix for at least one frequency band of the binaural audio signal; generate an energy estimate for at least one frequency band of the binaural audio signal; generate at least a part of a target covariance matrix for at least one frequency band of the binaural audio signal based on at least one direction parameter of at least one frequency band; generate a mixing matrix for at least one frequency band for mixing the binaural audio signal; and generate a left-channel audio signal and a right-channel audio signal from a combination of channels of at least one frequency band of the binaural audio signal based on the mixing matrix.
[0047] At least two audio signals for speaker reproduction may include a left-channel audio signal and a right-channel audio signal.
[0048] The apparatus, which is caused to process a binaural audio signal to generate at least two audio signals for loudspeaker reproduction by modifying the inter-channel differences of at least one frequency band based on at least one direction parameter of at least one frequency band, may further be caused to: generate a decorrelated audio signal for at least one frequency band based on the binaural audio signal; generate another mixing matrix for the decorrelated audio signal; generate another left-channel audio signal and another right-channel audio signal from channel combinations of at least one frequency band of the decorrelated audio signal based on the another mixing matrix; combine the left-channel audio signal and the another left-channel audio signal to generate a combined left channel; and combine the right-channel audio signal and the another right-channel audio signal to generate a combined right channel, and wherein the at least two audio signals for loudspeaker reproduction include the combined left-channel audio signal and the combined right-channel audio signal.
[0049] The apparatus, which is caused to further apply spectral adjustment to the processed at least one frequency band based on at least one direction parameter of at least one frequency band, may be caused to: determine a binaural response and / or a long-term response estimate based on the direction parameter of at least one frequency band; and compensate for the determined binaural response and / or long-term response estimate from the processed at least one frequency band.
[0050] The binaural response and / or the long-term response may include at least one of the following: at least one energy / amplitude; at least one correlation of channels of the binaural audio signal; at least one phase difference of channels of the binaural audio signal; at least one time difference of channels of the binaural audio signal.
[0051] The binaural response and / or the long-term response may include the spectrum of the binaural audio signal, and wherein the apparatus, which is caused to remove the determined binaural response and / or long-term response estimate from the processed at least one frequency band, may be caused to: obtain a filter and / or a gain based on the estimated direction parameter and an average head-related transfer function corresponding to the at least one direction parameter; and apply the filter and / or the gain to the processed at least one frequency band.
[0052] The apparatus, which is caused to determine a binaural response and / or a long-term response estimate based on the direction parameter of at least one frequency band, may be caused to: generate a long-term equalization filter by comparing an average spectrum of the binaural signals with a predetermined HRTF data set, and wherein the apparatus, which is caused to remove the determined binaural response and / or long-term response estimate from the processed at least one frequency band, may be caused to: apply the long-term equalization filter to the processed at least one frequency band.
[0053] The apparatus, which is caused to obtain at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal, may be caused to: analyze at least one frequency band of the binaural audio signal to determine at least one direction parameter of at least one frequency band.
[0054] The apparatus configured to analyze at least one frequency band of a binaural audio signal to determine at least one direction parameter of the at least one frequency band may further be configured to: for the at least one frequency band, estimate a delay that maximizes the correlation between channels of the binaural audio signal; and based on the estimated delay, formulate a direction parameter.
[0055] The apparatus may further be configured to: for at least one frequency band of the binaural audio signal, obtain a direct-to-total energy ratio based on the measured normalized correlation between channels of the binaural audio signal.
[0056] The apparatus configured to generate at least a portion of a target covariance matrix for at least one frequency band of a binaural audio signal based on at least one direction parameter of the at least one frequency band may further be configured to: for at least one frequency band of the binaural audio signal, further generate at least a portion of the target covariance matrix based on the direct-to-total energy ratio of the at least one frequency band.
[0057] The apparatus configured to determine a binaural response and / or a long-term response estimate based on at least one direction parameter of at least one frequency band may be configured to: determine the binaural response and / or the long-term response estimate based on the direct-to-total energy ratio of the at least one frequency band.
[0058] The apparatus configured to obtain a binaural audio signal may be configured to perform one of the following: capture the binaural audio signal with an artificial head; capture the binaural audio signal at the entrance of a user's ear canal; render the binaural audio signal according to a head-related transfer function; and render the binaural audio signal using a binaural room impulse response.
[0059] The apparatus configured to output at least two audio signals for speaker reproduction may be configured to: output the at least two audio signals for speaker reproduction to a stereo speaker.
[0060] According to a fourth aspect, there is provided an apparatus, comprising: an acquisition circuit configured to acquire a binaural audio signal; an acquisition circuit configured to obtain at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal; a processing circuit configured to process the binaural audio signal to generate at least two audio signals for speaker reproduction by modifying an inter-channel difference of at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band; and an output circuit configured to output the at least two audio signals for speaker reproduction.
[0061] According to a fifth aspect, there is provided a computer program including instructions [or a computer-readable medium including program instructions] for causing a device to at least perform the following operations: obtaining a binaural audio signal; obtaining at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal; processing the binaural audio signal to generate at least two audio signals for speaker reproduction by modifying an inter-channel difference of at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band; and outputting the at least two audio signals for speaker reproduction.
[0062] According to a sixth aspect, there is provided a non-transitory computer-readable medium including program instructions for causing a device to at least perform the following operations: obtaining a binaural audio signal; obtaining at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal; processing the binaural audio signal to generate at least two audio signals for speaker reproduction by modifying an inter-channel difference of at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band; and outputting the at least two audio signals for speaker reproduction.
[0063] According to a seventh aspect, there is provided a device including: means for obtaining a binaural audio signal; means for obtaining at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal; means for processing the binaural audio signal to generate at least two audio signals for speaker reproduction by modifying an inter-channel difference of at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band; and means for outputting the at least two audio signals for speaker reproduction.
[0064] According to an eighth aspect, there is provided a computer-readable medium including program instructions for causing a device to at least perform the following operations: obtaining a binaural audio signal; obtaining at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal; processing the binaural audio signal to generate at least two audio signals for speaker reproduction by modifying an inter-channel difference of at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band; and outputting the at least two audio signals for speaker reproduction.
[0065] A device includes means for performing the actions of the method as described above.
[0066] A device is configured to perform the actions of the method as described above.
[0067] A computer program includes program instructions for causing a computer to perform the method as described above.
[0068] A computer program product stored on a medium can cause a device to perform the methods described herein.
[0069] An electronic device can include a device as described herein.
[0070] A chipset can include a device as described herein.
[0071] Embodiments of the present application are intended to solve problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] To better understand the present application, reference will now be made, by way of example, to the accompanying drawings, in which:
[0073] Figure 1 A system schematically showing a device suitable for implementing some embodiments;
[0074] Figure 2 A flowchart showing the operation of an example device according to some embodiments;
[0075] Figure 3 Schematically showing, according to some embodiments, such as Figure 1 the inter-channel difference modifier shown in;
[0076] Figure 4 Showing, according to some embodiments, such as Figure 3 a flowchart of the operation of an example inter-channel difference modifier shown in;
[0077] Figure 5 Schematically showing, according to some embodiments, such as Figure 1 the spectral whitener shown in;
[0078] Figure 6 Showing, according to some embodiments, such as Figure 5 a flowchart of the operation of an example spectral whitener shown in; and
[0079] Figure 7 Showing an example device suitable for implementing the devices shown in the previous drawings. DETAILED DESCRIPTION
[0080] Suitable devices and possible mechanisms for converting binaural signals into conventional stereo audio signals are described in more detail below.
[0081] The concept, discussed in further detail in the following embodiments, is to generate a suitable stereo audio signal from a binaural audio signal. In the following description, at least two audio signals are generated, which may include left and right channel audio signals, or may include front, center, rear, top, or bottom versions of the left and right channels. The generated stereo audio signal can be reproduced using (stereo) speakers. Thus, the binaural cues (ITD, ILD, IC, spectral cues) generated at the listener's ears by the reproduction of the generated stereo audio signal by the stereo speakers are similar to the binaural cues when the binaural signal is played back on headphones, and spatial audio is perceived in the expected manner. In other words, it aims to prevent the perceptual differences caused at the listener's ears depending on the output component. These differences may include: differences in sound direction, differences in sound width, differences in sound spatial sense, differences in the spectrum of the sound.
[0082] Regarding spectral differences, binaural signals typically contain a unique spectrum caused by reflections from the human ears, head, torso, etc. The embodiments discussed herein are aimed at generating a stereo audio signal based on the binaural signal, wherein this unique spectrum is compensated so that when reproduced using (stereo) speakers and listened to by a human listener, there is no additional binaural response at the signal. Thus, the human listener does not receive "double binaural spectra", and the perception of timbre is similar to the original timbre.
[0083] Regarding direction differences, binaural signals are close to being actually two mono signals with potential phase differences at lower frequencies. Therefore, reproducing such a signal using stereo speakers produces an effect similar to shifting the sound amplitude to the middle of the speaker pair at lower frequencies. The embodiments discussed herein attempt to generate a stereo audio signal that maintains a proper perception of width and source localization when reproduced using a stereo speaker configuration compared to the binaural audio signal when reproduced using headphones.
[0084] The embodiments discussed herein are configured to generate a suitable stereo audio signal from a binaural audio signal, thus preventing the need to use binaural signals when using stereo speakers as the playback component, thereby preventing or reducing any spatial and timbre perception errors. Thus, the embodiments discussed herein have improved perceptual audio quality because the sound source is not perceived from the wrong direction and the timbre is not colored by the binaural audio signal directly reproduced using stereo speakers.
[0085] The concepts discussed in the embodiments herein can be generalized to apparatuses and methods related to reproducing binaural signals with loudspeakers, and in which apparatuses and / or methods are provided for converting binaural signals into "non-binaural" stereo signals suitable for stereo and multi-channel loudspeaker reproduction. Additionally, as described in the embodiments herein, the conversion is performed by analyzing the direction of arrival (or more generally, direction parameters) in the frequency bands of the binaural signals and modifying the binaural signals based on the analyzed directions such that the inter-channel differences and spectra match the expected characteristics of the "non-binaural" stereo signals.
[0086] The binaural signals can be of any kind of binaural signals, such as: signals captured with an artificial head, signals captured at the entrance of the ear canals of a real human, signals rendered using head-related transfer functions, or signals rendered using binaural room impulse responses. Additionally, the binaural signals may or may not contain any type of headphone compensation (e.g., derived using the measured headphone transfer function).
[0087] Binaural signals are intended for headphone listening, and when doing so, they create a natural perception of spatial sound (via natural ITDs, ILDs, and spectra). Thus, the sound source can be perceived from the correct direction with the correct timbre. In contrast, "non-binaural" stereo signals are intended for loudspeaker listening (i.e., they are "conventional" stereo signals). If listened to through headphones, the reproduction does not resemble binaural sound in terms of ITDs, ILDs, nor the binaural spectrum, but instead, these characteristics are formed when the "non-binaural" stereo signal is reproduced by a loudspeaker and propagated to the listener's ears.
[0088] The direction of arrival can be analyzed by estimating the delay that maximizes the correlation between the (binaural) signals in the frequency bands and formulating the direction value based on that delay value. Based on the measured normalized correlation between the binaural left and right signals, the direct-to-total energy ratio is estimated in the frequency bands.
[0089] In some embodiments, the inter-channel differences can be modified by at least determining the target energy / amplitude (and potentially correlation, phase / time difference) for loudspeaker reproduction based on the direction and ratio metadata and at least correcting the energy / amplitude (and potentially correlation, phase / time difference) of the input binaural signals to match the corresponding target characteristics.
[0090] In some embodiments, the spectrum can be modified by first obtaining a filter (or gain in the frequency band) based on the estimated direction of arrival and the average HRTF corresponding to that direction (of the set of HRTFs). Additionally, a long-term equalization filter can be applied by comparing the average spectrum of the binaural signals with a predetermined HRTF data set (also with different headphone compensations).
[0091] In some embodiments, the resulting "non-binaural" signals substantially remove or reduce any binaural characteristics (inherent in the original binaural signals) in them. Thus, the binaural characteristics will be added through the sound propagation from the speakers to the listener's ears. Therefore, for the speaker reproduction of the binaural signals using the present invention, good audio quality (accurate and natural direction perception and non-colored timbre) can be achieved.
[0092] Regarding Figure 1 , a block diagram of a device suitable for implementing some embodiments is shown. As will be described subsequently, the device can be implemented inside a mobile phone or a computer. Additionally, it can be implemented, for example, as an independent device or program, or it can be part of an audio codec such as an IVAS codec, for example.
[0093] The block diagram shows a binaural audio signal 100. In this example, the binaural audio signal 100 is a time-domain signal. However, in some embodiments where the binaural audio signal 100 is a time-frequency domain signal, the use of the time-frequency transformer can be skipped or bypassed.
[0094] In some embodiments, the device includes a time-frequency transformer 101. The time-frequency transformer 101 is configured to receive (time-domain) binaural audio signals 100 and convert them to the time-frequency domain. Suitable transforms include, for example, the short-time Fourier transform (STFT) and the polyphase modulated quadrature mirror filter (QMF) bank. The resulting time-frequency binaural audio signals 102 can be labeled as S m (b,n), where m is the channel index, b is the frequency bin index, and n is the time index.
[0095] The time-frequency binaural audio signals 102 can be forwarded to the direction analyzer 105 and the inter-channel difference modifier 103.
[0096] In some embodiments, the device or converter includes a direction analyzer 105. The direction analyzer 105 is configured to receive the time-frequency domain binaural audio signals 102 and analyze the direction of arrival θ(k,n) and the direct-to-total energy ratio r(k,n) in the time-frequency domain, where k is the frequency band index.
[0097] The direction analysis is performed in the frequency bands. The time-frequency transform has a certain frequency resolution. For example, a 1024-point STFT results in 513 frequency bins from the DC frequency to the Nyquist frequency. These bins are grouped into frequency bands, for example, 24 frequency bands with a resolution close to the Bark frequency resolution.
[0098] The analysis can be performed within these frequency bands. Each frequency band k has the lowest bin b low (k) and the highest bin b high (k).
[0099] The analyzer can be configured, for example, to find the delay τ k which maximizes the correlation between the two channels for each frequency band k. This can be achieved by creating time-shifted versions of the signal in one of the channels and correlating these time-shifted versions with the other channel signal. S m The time shift of τ time-domain samples of s(b,n) can be obtained as:
[0100]
[0101] where N is the length of the STFT operation. The optimal delay τ for frequency band k (and time index n) k is obtained from:
[0102]
[0103] where c(k,n) is the correlation with the optimal delay τ k (which is the parameter τ that maximizes the above equation), Re denotes the real part of the result, and * denotes the complex conjugate. The search range for the delay D max is selected based on the estimated maximum time delay difference between the sound arriving at the two ears.
[0104] The delay τ k can be transformed into an angular value by:
[0105]
[0106] This direction parameter is an azimuthal value between -90 and 90 degrees. This direction information 106 is sufficient for rendering the stereo speaker output, since there are no elevated or rear speakers (in other words, the output audio signal is in the "horizontal" plane and no elevation value is required). The direction information 106 or the signal can then be output to the inter-channel modifier 103 and the spectral whitener 107.
[0107] Additionally, in some embodiments, the direction analyzer 105 is further configured to determine at least one corresponding energy ratio r(k,n). The energy ratio r(k,n) can be estimated by the following operation: using, for example, the normalized correlation value c(k,n), e.g., by:
[0108]
[0109] and then comparing the correlation value with the inter-aural diffuse-field correlation at the center frequency of the frequency band c diff (k) to obtain the ratio:
[0110]
[0111] The estimated direct-to-total energy ratio can also be forwarded to the inter-channel difference modifier 103 and the spectral whitener 107.
[0112] In some embodiments, the converter includes an inter-channel difference modifier 103. The inter-channel difference modifier 103 is configured to receive the time-frequency binaural audio signal 102, the direction information 106, and the energy ratio information 108. The inter-channel difference modifier 103 is configured to modify, based on the analyzed direction and energy ratio, at least the inter-aural level difference (and potentially the phase and / or time difference and / or coherence) of the time-frequency binaural audio signal in the frequency band(s) such that the processed output has an inter-channel level difference suitable for speaker reproduction of sound in the direction θ(k,n) with a direct-to-total energy ratio r(k,n).
[0113] The resulting time-frequency intermediate audio signal 104 is output from the inter-channel difference modifier and passed to the spectral whitener 107.
[0114] In some embodiments, the converter includes a spectral whitener 107. The spectral whitener is configured to receive the time-frequency intermediate audio signal 104. The time-frequency intermediate audio signal 104 has direction cues (e.g., level differences) suitable for speaker playback, but they still have elements of the binaural spectrum included therein, which can be removed using the spectral whitener 107. Accordingly, the spectral whitener 107 is further configured to receive the direction information 106 and the direct-to-total energy ratio information 108. The spectral whitener 107 is configured to invert or compensate the binaural spectrum, and the resulting time-frequency stereo audio signal 110 is output to the inverse time-frequency transformer 111.
[0115] In some embodiments, the converter includes an inverse time-frequency transformer 111. The inverse time-frequency transformer 111 is configured to apply an inverse transform corresponding to the applied time-frequency transform (e.g., an inverse STFT corresponding to the STFT) to the received time-frequency stereo audio signal 110 and is configured to output a suitable (pulse-code modulated) PCM stereo audio signal 112, which can then be reproduced with stereo speakers.
[0116] Regarding Figure 2 ,, a flowchart illustrating the operation of the converter as shown in Figure 1 is shown.
[0117] Thus, for example, the first operation is the operation of receiving the binaural audio signal as shown by step 201 in Figure 2 .
[0118] Furthermore, as shown by step 203 in Figure 2 , a time-frequency transform is applied to the binaural audio signal to generate a time-frequency binaural audio signal.
[0119] As inFigure 2 As shown in step 204, the time-frequency binaural audio signal can then be analyzed to determine the direction and energy ratio.
[0120] As in Figure 2 As shown in step 205, the time-frequency binaural audio signal can then be modified between channels based on the determined direction and energy ratio to generate a time-frequency intermediate audio signal.
[0121] As in Figure 2 As shown in step 207, the time-frequency intermediate audio signal can also be spectrally whitened based on the determined direction and energy ratio to generate a time-frequency processed (stereo) audio signal.
[0122] Furthermore, as in Figure 2 As shown in step 209, an inverse time-frequency transform is performed on the time-frequency processed (stereo) audio signal to generate a stereo audio signal.
[0123] As in Figure 2 As shown in step 211, the stereo audio signal can then be output.
[0124] Regarding Figure 3 , the inter-channel difference modifier 103 is shown in more detail. In some embodiments, the inter-channel difference modifier 103 includes a covariance matrix estimator 301. The covariance matrix estimator 301 is configured to receive the time-frequency binaural audio signal 102 and generate a suitable estimated covariance matrix (Estimated cov mtx) 300, such as:
[0125]
[0126] where H represents the complex conjugate, and
[0127]
[0128] The covariance matrix estimator 301 is configured to output the estimated covariance matrix C in (k,n) 300 to the mixing matrix formulator 307.
[0129] The covariance matrix estimator 301 can also be configured to formulate the total energy estimate E(k,n) as the sum of the diagonal elements of C in (k,n). The total energy estimate 302 is provided to the target covariance matrix formulator 305.
[0130] In the examples described herein, the input and target covariance matrix formulas encapsulate a set of inter-channel characteristics (energy difference, phase difference, correlation), and all of these can be processed. However, in some embodiments, there may be at least a portion of the signal (e.g., at some frequencies) where only the energy will be adjusted or modified. In such cases, it is not necessary to estimate the full covariance matrix. However, for simplicity, the full covariance matrix is estimated herein, and then data that may not be necessary (depending on the configuration) is not used at a later stage. In some embodiments, the actual implementation is configured to estimate only the data or information required at a later stage.
[0131] In some embodiments, the inter-channel difference modifier 103 includes a target covariance matrix formulator 305. The target covariance matrix formulator 305 is configured to receive the energy estimate 302 and the direction θ(k,n) and the direct-to-total energy ratio r(k,n) parameters 108. In some embodiments, the target covariance matrix formulator 305 generates a target covariance matrix for the output speaker signals. In some embodiments, this can be achieved by the following operations.
[0132] First, the matrix generates a translation gain:
[0133]
[0134] where g L (θ(k,n)) and g R (θ(k,n)) are the gains according to the Vector Base Amplitude Translation (VBAP) law for speakers at ±30°:
[0135]
[0136] and
[0137]
[0138] Furthermore, the target covariance matrix is formulated as:
[0139]
[0140] where the left part g(k,n)g T (k,n)r(k,n) represents the covariance matrix associated with the previously translated sound, and the right part represents the covariance matrix associated with the ambient (or non-directional) sound.
[0141] As shown in the above equations, these left and right parts are then added together and weighted by the total energy estimate E(k,n) to obtain the target covariance matrix C target (k,n).
[0142] Furthermore, the target covariance matrix C target (k,n)306 can be provided to the mixing matrix formulator 307.
[0143] In some embodiments, the inter-channel difference modifier 103 includes the mixing matrix formulator 307. The mixing matrix formulator 307 is configured to receive the target covariance matrix 306 and the estimated covariance matrix 300 and generate a mixing matrix 308, which can be passed to the mixer 309.
[0144] In some embodiments, the mixing matrix formulator 307 is configured to generate the mixing matrix according to the method described in US20140233762A1 and "Optimized covariance domain framework for time-frequency processing of spatial audio" by Vilkamo Juha, Tom and Achim Kuntz (Journal of the Audio Engineering Society, Vol. 61, No. 6, 2013: pp. 403-411).
[0145] The methods in the cited papers include least squares optimization signal mixing techniques for manipulating the covariance matrix of signals while well maintaining the audio quality. Therefore, these methods utilize the covariance matrix metrics of the input signals and the target covariance matrix and provide a mixing matrix to perform such processing. When there is not enough independent signal energy at the input end, these methods also provide means to optimally use decorrelated sounds.
[0146] Therefore, in some embodiments, the mixing matrix formulator 307 is configured to generate a prototype matrix that determines how the output channels should resemble the input channels (while satisfying the synthesis of the target covariance matrix). In the current context, the prototype matrix is:
[0147]
[0148] When Q, C target (k,n) and C in (k,n) are now known, the methods discussed in the cited papers provide two mixing matrices M(k,n) for non-decorrelated sounds and M r (k,n) for decorrelated sounds. These mixing matrices 308 are provided to the mixer 309.
[0149] In some embodiments, the mixing matrix formulator 307 is configured to compensate (only) for the energy of the signals without affecting the phase or correlation between channels. For example, at high frequencies, this can be the most robust option, and at high frequencies, the phase / correlation information also has less perceptual relevance than at low frequencies. In this case, the mixing matrix expressed by the formula can be:
[0150]
[0151] where the braces {} denote selecting a single matrix entry from the covariance matrix. The rest of the processing is as described above.
[0152] In some embodiments, the inter-channel difference modifier 103 includes a channel decorrelator 303. The channel decorrelator 303 is configured to receive the time-frequency binaural audio signal 102 and apply decorrelation to the two channels s(b,n) to generate two uncorrelated versions of the binaural input signal (relative to each other and relative to the input). The result is the decorrelated signal s d (b,n). The decorrelation process can be a time-invariant phase-scrambling process. Any decorrelator can be applied, and the choice of decorrelator can depend on the time-frequency transform applied. The decorrelated signal 304 is then provided to the mixer 309.
[0153] In some embodiments, the inter-channel difference modifier 103 includes a mixer 309. The mixer 309 is configured to receive the time-frequency decorrelated audio signal 304, the time-frequency binaural audio signal 102, and the mixing matrix 308, and generate the time-frequency speaker signal 104 (without spectral whitening) for each frequency band k as:
[0154] s′ LS (b,n) = M(k,n)s(b,n) + M r (k,n)s d (b,n)
[0155] The mixing matrix is for each frequency band k, and the same mixing matrix can be applied for each bin b within that frequency band. The mixing matrix (or alternatively, the covariance matrix before formulating the mixing matrix) can be smoothed over time to reduce potential processing artifacts. The mixer 309 is then configured to output the time-frequency intermediate (speaker) signal (without spectral whitening) 104.
[0156] The operation of the inter-channel difference modifier 103 is shown in the flowchart as shown in Figure 4 .
[0157] Receiving the time-frequency binaural audio signal is atFigure 4 is shown by step 401 in
[0158] As in Figure 4 shown by step 403 in
[0159] Furthermore, as in Figure 4 shown by step 405 in
[0160] Receiving parametric parameters such as direction and energy ratio is shown by step 404 in Figure 4
[0161] As in Figure 4 shown by step 407 in
[0162] As in Figure 4 shown by step 409 in
[0163] Furthermore, as in Figure 4 shown by step 411 in
[0164] As in Figure 4 shown by step 413 in
[0165] Regarding Figure 5 is shown a block diagram of an example spectral whitener 107 according to some embodiments.
[0166] The spectral whitener 107 is configured to receive a time-frequency intermediate (speaker) signal (without spectral whitening) s′ LS (b,n)104, a direction θ(k,n)106, and a direct-to-total energy ratio r(k,n)108.
[0167] In some embodiments, the spectral whitener 107 includes a binaural response estimator 503. In some embodiments, the binaural response estimator 503 is configured to receive the direction 106 and the energy ratio 108. The binaural response estimator 503 may then estimate the energy response of a typical binaural signal corresponding to the direction θ(k,n) and the energy ratio r(k,n). This energy response is common to both ears because the inter-channel differences have been corrected in the inter-channel difference modifier 103.
[0168] The binaural response estimator 503 can be configured, for example, to first estimate the energy response to the direct sound based on the direction θ(k,n). This can be achieved, for example, by the following formula:
[0169] E dir (k,n) = f HRTF (θ(k,n))
[0170] where f HRTF () is a function for obtaining the average energy spectrum of the HRTF pair corresponding to the direction θ in the frequency band k. It can be implemented in any suitable way. For example, several sets of HRTFs are obtained. In this example, each set of HRTFs has the same set of directions in the dataset. Next, the average energy response of the HRTF pair is calculated for each direction in each dataset. For example, by the following formula:
[0171]
[0172] where H left is the HRTF for the left ear, H right is the HRTF for the right ear, i is the index of the dataset, and |.| denotes calculating the absolute value. When the HRTF in the frequency band k is determined, the HRTF at the mid-frequency of the frequency band k can then be formulated. The datasets can be combined, for example, by taking their average for each direction, so as to obtain E avg (k,θ). Then, finally, f HRTF () can be implemented, for example, by interpolating between the closest data points of E avg (k,θ) to obtain the value for the direction θ (if the dataset E avg (k,θ) has a data point exactly at the direction θ, it can be used directly).
[0173] Next, the energy response to the ambient sound is estimated. Since this estimate is not based on any parameters, it can be obtained from the database. The estimate of the ambient sound energy response can be formed, for example, by averaging over all directions of the averaged HRTF energy dataset:
[0174]
[0175] where θ(d) are the D HRTF directions in the dataset.
[0176] Then, the estimate of the binaural energy response can be formed by the following formula:
[0177] E bin (k,n) = r(k,n)E dir (k,n) + (1 - r(k,n))E amb (k)
[0178] It can be output as a binaural response 504 to the binaural response remover 501.
[0179] In some embodiments, the spectral whitener 107 includes a binaural response remover 501. The binaural response remover 501 is configured to receive the time-frequency intermediate (speaker) signal (without spectral whitening) s′ LS (b,n) 104 and the binaural energy response E bin (k,n) 504 as inputs. The binaural response remover 501 is configured to first formulate an equalizer by the following formula:
[0180]
[0181] which can be smoothed in time (alternatively, E bin (k,n) can be smoothed in time before formulating g EQ (k,n)). Further, a set of processed intermediate signals can be expressed by the formula, such as by the following formula:
[0182] S″ LS (b,n) = g EQ (k,n) s′ LS (b,n)
[0183] where k is the frequency band index where bin b is located. At the obtained processed intermediate signal s″ LS (b,n), the binaural spectrum according to the average HRTF has been removed. Generally, these signals are already suitable for speaker reproduction. However, due to possible differences in the way binaural signals are initially generated (e.g., there are different types of artificial heads and different HRTF and BRIR databases), the spectrum of the processed intermediate signal s″ LS (b,n) 502 may still deviate from the optimal value.
[0184] Therefore, in some embodiments, the processed intermediate signal s″ LS (b,n) 502 can be forwarded to the long-term response estimator 505 (which is also referred to as the long-term spectrum estimator) and the long-term response remover 507.
[0185] In some embodiments, the spectral whitener 107 includes a long-term response estimator 505, which is configured to receive the processed intermediate signal s″ LS (b,n) 502, estimate the long-term spectrum of these intermediate signals and compare it with the expected average spectrum. If the estimator finds a reliable deviation between the two, it generates the estimated long-term response H lt (b,n) 506 and sends it to the long-term response remover 507.
[0186] In some embodiments, the spectral whitener 107 includes a long-term response remover 507 configured to receive the processed intermediate signal s″ LS (b,n)502 and process the processed intermediate signal s″ LS (b,n)502 based on the estimated long-term response 506, and output a suitable time-frequency stereo (loudspeaker) audio signal 110:
[0187]
[0188] When no reliable deviation is detected, the estimated response H lt (b,n) can be set to 1 at all frequencies. Additionally, in some embodiments, the long-term response estimator 505 and the long-term response remover 507 are optional and can be omitted, and the processed intermediate audio signal s″ LS (b,n)502 is directly passed as the time-frequency stereo audio signal 110.
[0189] The output of the spectral whitener 107 is shown as the time-frequency domain stereo signal s LS (b,n), which is then transformed into a time-domain signal as expressed in the context of Figure 1 and the result is suitable for loudspeaker reproduction.
[0190] The inter-aural channel differences have been modified to inter-aural channel differences more suitable for loudspeaker reproduction, and the binaural spectra have been compensated.
[0191] Regarding Figure 6 , a flowchart illustrating the operation of the exemplary spectral whitener 107 is shown.
[0192] Thus, as shown by step 601 in Figure 6 , the time-frequency intermediate audio signal is received.
[0193] Additionally, the parametric parameters such as direction and energy ratio are received as shown by step 602 in Figure 6 .
[0194] As shown by step 604 in Figure 6 , the binaural response is estimated.
[0195] Subsequently, as shown by step 605 in Figure 6 , the estimated binaural response is removed from the time-frequency intermediate audio signal.
[0196] Optionally, as shown by step 607 in Figure 6 , the long-term response is further estimated.
[0197] Subsequently, as shown by step 608 in Figure 6As shown in step 609, optionally, the estimated long-term response is further removed.
[0198] In the embodiments discussed above, the binaural signals are fully converted into non-binaural stereo signals. However, there may be cases where it is desired to convert only a part of the binaural signals into non-binaural stereo signals. For example, when the binaural-to-non-binaural conversion occurs, only those directions mapped between the stereo speakers can be rendered as non-binaural sounds, while using a cross-talk cancelling scheme to reproduce the remaining (binaural) sounds on the speakers. Thus, in some embodiments, a part of the binaural audio signal for a range of directions is converted into a stereo signal, while the remaining part of the signal is passed through without conversion. This part can also be a part of the total energy of the binaural audio signal, or can be a part of the spectrum of the binaural audio signal (e.g., some of the frequency bands are converted while some of the frequency bands are passed through without being processed).
[0199] Regarding Figure 7 , an example electronic device that can be used as any device component of the system described above is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1700 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0200] In some embodiments, device 1700 includes at least one processor or central processing unit 1707. The processor 1707 can be configured to execute various program codes, such as the methods described herein.
[0201] In some embodiments, device 1700 includes a memory 1711. In some embodiments, at least one processor 1707 is coupled to the memory 1711. The memory 1711 can be any suitable storage component. In some embodiments, the memory 1711 includes a program code portion for storing program codes that can be implemented on the processor 1707. Additionally, in some embodiments, the memory 1711 can also include a stored data portion for storing data (e.g., data that has been processed or will be processed according to the embodiments described herein). As long as needed, the implemented program codes stored in the program code portion and the data stored in the stored data portion can be obtained by the processor 1707 via the memory-processor coupling.
[0202] In some embodiments, device 1700 includes user interface 1705. In some embodiments, user interface 1705 may be coupled to processor 1707. In some embodiments, processor 1707 may control the operation of user interface 1705 and receive input from user interface 1705. In some embodiments, user interface 1705 may enable a user to input commands to device 1700, for example, via a keypad. In some embodiments, user interface 1705 may enable a user to obtain information from device 1700. For example, user interface 1705 may include a display configured to display information from device 1700 to the user. In some embodiments, user interface 1705 may include a touch screen or touch interface that can both enable information to be input into device 1700 and display information to the user of device 1700. In some embodiments, user interface 1705 may be a user interface for communication.
[0203] In some embodiments, device 1700 includes input / output port 1709. In some embodiments, input / output port 1709 includes a transceiver. In such embodiments, the transceiver may be coupled to processor 1707 and configured to communicate with other devices or electronic equipment, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components may be configured to communicate with other electronic devices or equipment via a wired or wired coupling.
[0204] The transceiver may communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication path (IRDA).
[0205] The transceiver input / output port 1709 may be configured to receive signals.
[0206] Input / output port 1709 may be coupled to any suitable audio output, for example, a stereo speaker system.
[0207] Generally, various embodiments of the present invention can be implemented using hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented using hardware, while other aspects can be implemented using firmware or software executable by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.
[0208] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Additionally, in this regard, it should be noted that any block of the logical flow in the drawings can represent a program step, or interconnected logical circuits, blocks, and functions, or a combination of program steps and logical circuits, blocks, and functions. The software can be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or a floppy disk, and on an optical medium such as a DVD and its data variant CD.
[0209] The memory can be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, as a non-limiting example, can include a general-purpose computer, a dedicated computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and one or more of the processors.
[0210] Embodiments of the present invention can be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools can be used to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0211] Programs, such as those provided by Synopsys, Inc., of Mountain View, California and Cadence Design of San Jose, California, use well-established design rules and pre-stored libraries of design modules to automatically route conductors and place components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transferred to a semiconductor manufacturing facility or "fab" for fabrication.
[0212] The foregoing description has provided a complete and useful description of exemplary embodiments of the invention by way of example and not limitation. However, various modifications and adaptations will become apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the invention will still fall within the scope of the invention as defined by the appended claims.
Claims
1. An apparatus for converting a binaural signal into a stereo audio signal, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: Obtain a binaural audio signal; Based on an analysis of at least one frequency band of the binaural audio signal, obtain at least one direction parameter of the at least one frequency band of the binaural audio signal; Based on the at least one direction parameter of the at least one frequency band, process the binaural audio signal by modifying the inter-channel difference of the at least one frequency band of the binaural audio signal to generate at least two audio signals for loudspeaker reproduction; and Output the at least two audio signals for loudspeaker reproduction.
2. The device according to claim 1, wherein, The inter-channel difference includes at least one of the following: At least one energy / amplitude difference between channels of the binaural audio signal; At least one phase difference between channels of the binaural audio signal; and At least one time difference between channels of the binaural audio signal.
3. The device according to claim 1, wherein Process the binaural audio signal such that the apparatus further applies spectral adjustment to the processed at least one frequency band based on the at least one direction parameter of the at least one frequency band.
4. The device according to claim 1, wherein, Process the binaural audio signal such that the apparatus: Generate an estimate of at least a part of a covariance matrix for at least one frequency band of the binaural audio signal; Generate an energy estimate for at least one frequency band of the binaural audio signal; Generate at least a part of a target covariance matrix for at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band; Generate a mixing matrix for mixing at least one frequency band of the binaural audio signal; And Based on the mixing matrix, generate a left-channel audio signal and a right-channel audio signal from a channel combination of at least one frequency band of the binaural audio signal.
5. The device according to claim 4, wherein The at least two audio signals for loudspeaker reproduction include the left-channel audio signal and the right-channel audio signal.
6. The device according to claim 4, wherein Process the binaural audio signal such that the apparatus: Generate a decorrelated audio signal for at least one frequency band based on the binaural audio signal; Generate another mixing matrix for the decorrelated audio signal; Based on the another mixing matrix, generate another left-channel audio signal and another right-channel audio signal from a channel combination of at least one frequency band of the decorrelated audio signal; Combine the left-channel audio signal and the another left-channel audio signal to generate a combined left channel; And Combine the right-channel audio signal and the another right-channel audio signal to generate a combined right channel, and wherein the at least two audio signals for loudspeaker reproduction include the combined left-channel audio signal and the combined right-channel audio signal.
7. The device according to claim 3, wherein Cause the apparatus to further apply spectral adjustment to the processed at least one frequency band based on the at least one direction parameter of the at least one frequency band such that the apparatus: Determine a binaural response and / or a long-term response estimate based on the direction parameter of the at least one frequency band; and Compensate the determined binaural response and / or long-term response estimate from the at least one processed frequency band.
8. The apparatus according to claim 7, wherein The binaural response and / or the long-term response includes at least one of the following: At least one energy / amplitude; At least one correlation of the channels of the binaural audio signal; At least one phase difference of the channels of the binaural audio signal; At least one time difference of the channels of the binaural audio signal.
9. The device according to claim 7, wherein The binaural response and / or the long-term response includes the spectrum of the binaural audio signal, and wherein causing the device to remove the determined binaural response and / or long-term response estimate from the at least one processed frequency band causes the device to: Obtain a filter and / or gain based on the estimated direction parameter and the average head-related transfer function corresponding to the at least one direction parameter; and Apply the filter and / or the gain to the at least one processed frequency band.
10. The device according to claim 7, wherein, Causing the device to determine a binaural response and / or long-term response estimate based on the direction parameter of the at least one frequency band causes the device to: generate a long-term equalization filter by comparing the average spectrum of the binaural signals with a predetermined HRTF dataset, and wherein causing the device to remove the determined binaural response and / or long-term response estimate from the at least one processed frequency band further causes the device to: apply the long-term equalization filter to the at least one processed frequency band.
11. The device according to claim 1, wherein, Causing the device to obtain at least one direction parameter of at least one frequency band of the binaural audio signal based on the binaural audio signal further causes the device to: analyze the at least one frequency band of the binaural audio signal to determine the at least one direction parameter of the at least one frequency band.
12. The device according to claim 11, wherein, Causing the device to analyze the at least one frequency band of the binaural audio signal to determine the at least one direction parameter of the at least one frequency band further causes the device to: Estimate a delay for the at least one frequency band that maximizes the correlation between the channels of the binaural audio signal; and Formulate a direction parameter based on the estimated delay.
13. The device according to claim 1, wherein, Causing the device to: obtain a direct-to-total energy ratio for the at least one frequency band of the binaural audio signal based on the measured normalized correlation between the channels of the binaural audio signal.
14. The apparatus according to claim 13, wherein Causing the device to generate at least a portion of a target covariance matrix for the at least one frequency band of the binaural audio signal based on the at least one direction parameter of the at least one frequency band further causes the device to: generate at least a portion of the target covariance matrix for the at least one frequency band of the binaural audio signal further based on the direct-to-total energy ratio of the at least one frequency band.
15. The device according to claim 13, wherein, Causing the device to determine a binaural response and / or long-term response estimate based on the at least one direction parameter of the at least one frequency band causes the device to: determine the binaural response and / or long-term response estimate based on the direct-to-total energy ratio of the at least one frequency band.
16. The device according to claim 1, wherein, Causing the device to obtain a binaural audio signal causes the device to perform one of the following: Capture the binaural audio signal with a dummy head; Capture the binaural audio signals at the entrance of the user's ear canal; Render the binaural audio signals according to the head-related transfer function; And Render the binaural audio signals using the binaural room impulse response.
17. The device according to claim 1, wherein Cause the device to output the at least two audio signals for speaker reproduction, and further cause the device to: output the at least two audio signals for speaker reproduction to stereo speakers.
18. A method for converting binaural signals into stereo audio signals, comprising: Obtain binaural audio signals; Based on the analysis of at least one frequency band of the binaural audio signals, obtain at least one direction parameter of the at least one frequency band of the binaural audio signals; Based on the at least one direction parameter of the at least one frequency band, process the binaural audio signals by modifying the inter-channel differences of the at least one frequency band of the binaural audio signals to generate at least two audio signals for speaker reproduction; and Output the at least two audio signals for speaker reproduction.
19. The method according to claim 18, wherein, The inter-channel differences include at least one of the following: At least one energy / amplitude difference between channels of the binaural audio signals; At least one phase difference between channels of the binaural audio signals; and At least one time difference between channels of the binaural audio signals.
20. A non-transitory computer-readable medium, comprising program instructions for causing a device to at least perform the following operations: Obtain binaural audio signals; Based on the analysis of at least one frequency band of the binaural audio signals, obtain at least one direction parameter of the at least one frequency band of the binaural audio signals; Based on the at least one direction parameter of the at least one frequency band, process the binaural audio signals by modifying the inter-channel differences of the at least one frequency band of the binaural audio signals to generate at least two audio signals for speaker reproduction; and Output the at least two audio signals for speaker reproduction.
Citation Information
Patent Citations
Optimal mixing matrices and usage of decorrelators in spatial audio processing
US20140233762A1
Audio encoding and decoding
CN101390443A
Apparatus and method for stereo filling in multichannel coding
CN109074810A