Device and method for audio processing

By exporting and processing signal components depending on different parts of the spatial audio image, combining trans-channel mixing and headphone filters, the problem of sound quality degradation in stereo widening technology is solved, and the spatial audio performance in headphone devices is improved.

CN115190414BActive Publication Date: 2025-08-08NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210643129.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-29
Filing Date
2020-05-29
Publication Date
2025-08-08
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

When processing stereo widening signals, the existing stereo widening technology can easily lead to a decrease in the sound quality of the central part of the spatial audio image, especially in headphone equipment, which leads to a decrease in sound engagement and comb filtering effect, affecting the sound quality.

Method used

By deriving the first signal component and the second signal component depending on different parts of the spatial audio image, the second signal component is processed using a trans-channel mixing device, and bypassing the first signal component, combining the headphone filter and the head-dependent transformation function, an output audio signal suitable for rendering of the headphone device is generated.

Benefits of technology

It effectively maintains the sound quality of the central part of the spatial audio image, enhances the spatial audio perception in the headphone device, reduces the comb filtering effect, and improves the sound engagement and sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115190414B_ABST
    Figure CN115190414B_ABST
Patent Text Reader

Abstract

A device for processing an input audio signal comprising multiple channels, the device comprising: a device for deriving a first signal component comprising at least one input channel and a second signal component comprising multiple input channels based on the input audio signal, wherein the first signal component depends on at least a first part of a spatial audio image conveyed by the input audio signal and the second signal component depends on at least a second part of the spatial audio image that is different from the first part; a cross-channel mixing device for mixing the multiple input channels across channels; a device for guiding the second signal component to the cross-channel mixing device for cross-channel mixing of at least some of the multiple input channels of the second signal component to produce a modified second signal component; a bypass device for enabling the first signal component to bypass the cross-channel mixing device; and a device for combining the first signal component and the modified second signal component into an output audio signal, which includes two output channels configured for rendering by a headphone device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with application number 202010473489.X and invention name “Audio Processing” filed on May 29, 2020. Technical Field

[0002] Exemplary and non-limiting embodiments of the present invention relate to the processing of audio signals. In particular, various embodiments of the present invention relate to the modification of a spatial image represented by a multi-channel audio signal, such as a two-channel stereo signal. Background Art

[0003] Stereo widening is a technique known in the art for enhancing the perceived spatial audio image of a stereo audio signal when reproduced via an audio output device. Such techniques aim to process the stereo audio signal so that the reproduced sound is perceived not only as originating from directions positioned between the audio output devices, but also as if at least a portion of the sound field originates from directions not positioned between the audio output devices, thereby widening the perceived width of the spatial audio image conveyed in the stereo audio signal. This spatial audio image is referred to herein as a widened or enlarged spatial audio image.

[0004] Although summarized above with reference to a two-channel stereo audio signal, stereo widening can be applied to multi-channel audio signals having more than two channels, such as 5.1-channel or 7.1-channel surround sound for playback through a pair of audio output devices. In some contexts, the term virtual surround sound is applied to refer to processed audio signals that convey a spatial audio image originally conveyed in a multi-channel surround sound signal. Therefore, even though the term stereo widening is primarily used throughout this disclosure, the term should be interpreted broadly to cover techniques for processing a spatial audio image conveyed in a multi-channel audio signal (i.e., a two-channel stereo audio signal or surround sound of more than two channels) to provide audio playback over the widened spatial audio image.

[0005] For simplicity and clarity of description, in this disclosure, the term multi-channel audio signal is used to refer to an audio signal with two or more channels. In addition, the term stereo signal is used to refer to a stereo audio signal, and the term surround signal is used to refer to a multi-channel audio signal with more than two channels.

[0006] When applied to a stereo signal, stereo widening techniques known in the art generally involve adding a processed (e.g., filtered) version of a contralateral channel signal to each of the left and right channel signals of the stereo signal to derive an output stereo signal having a widened spatial audio image (hereinafter referred to as the widened stereo signal). In other words, a processed version of the right channel signal of the stereo signal is added to the left channel signal of the stereo signal to create the left channel of the widened stereo signal, and a processed version of the left channel signal of the stereo signal is added to the right channel signal of the stereo signal to create the right channel of the widened stereo signal. Furthermore, deriving the widened stereo signal may further involve pre-filtering (or otherwise processing) the corresponding processed contralateral signal before adding it to each of the left and right channel signals of the stereo signal to preserve a desired frequency response in the widened stereo signal.

[0007] Along the above ideas, stereo widening can be easily summarized as widening the spatial audio image of a multi-channel input audio signal, thereby deriving an output multi-channel audio signal with a widened spatial audio image (hereinafter referred to as the widened multi-channel signal). In this regard, the processing involves creating the left channel of the widened multi-channel audio signal as the sum of the (first) filtered versions of the channels of the multi-channel input audio signal, and creating the right channel of the widened multi-channel audio signal as the sum of the (second) filtered versions of the channels of the multi-channel input audio signal. Here, a dedicated predefined filter can be provided for each pair of input channels (channels of the multi-channel input signal) and output channels (left and right). As an example in this regard, the left and right channel signals of the widened multi-channel signal can be defined based on the channels of the multi-channel audio signal S, respectively, according to equation (1), S out,left and S out,right :

[0008]

[0009] Where S(i, b, n) represents the frequency bin b in the time frame n of the channel i of the multi-channel signal S, H left (i, b) represents the frequency bin b for filtering channel i of the multi-channel signal S to create the left channel signal S out,left (b, n) of the corresponding channel components, and H right (i, b) represents the frequency bin b for filtering channel i of the multi-channel signal S to create the right channel signal S out,right Filters for the corresponding channel components of (b, n).

[0010] A challenge involved in stereo widening is the degradation of the sound quality of the central part of the spatial audio image. In many real-life stereo signals, the central part of the spatial audio image includes perceptually important audio content, for example in the case of music, the singer's voice is usually rendered in the center of the spatial audio image. The sound component in the center of the spatial audio image is rendered by reproducing the same signal in both channels of the stereo signal and therefore via two audio output devices. When stereo widening is applied to such an input stereo signal (for example, according to equation (1) above), each channel of the resulting widened stereo signal involves the result of two filtering operations performed on the channels of the input stereo signal. This may result in a comb filtering effect, which can lead to differences in the perceived sound quality, which may be referred to as "coloration" of the sound. Moreover, the comb filtering effect may further lead to a degradation in the cohesion of the sound source.

[0011] In some cases, the audio output device is part of a headphone device that includes a left audio output device worn at, on, or in a user's left ear and a right audio output device worn at, on, or in a user's right ear.

[0012] Normal playback of stereo audio through headphones may cause the sound to be perceived by the user inside the user's head. Stereo panning suggests localizing the sound between the two ears inside the head.

[0013] To address this issue, speaker virtualization methods are used to process audio signals so that the user's experience of listening through headphones is similar to that of listening through speakers. This can be achieved by filtering the audio signal using an appropriate head-related transfer function (HRTF) or binaural room impulse response (BRIR). Summary of the Invention

[0014] According to various but not all examples, a device for processing an input audio signal comprising multiple channels is provided, the device comprising: a device for deriving, based on the input audio signal, a first signal component comprising at least one input channel and a second signal component comprising multiple input channels, wherein the first signal component depends on at least a first part of a spatial audio image conveyed by the input audio signal and the second signal component depends on at least a second part of the spatial audio image that is different from the first part; a cross-channel mixing device for mixing the multiple input channels across channels; a device for directing the second signal component to the cross-channel mixing device for cross-channel mixing of at least some of the multiple input channels of the second signal component to produce a modified second signal component; a bypass device for enabling the first signal component to bypass the cross-channel mixing device; and a device for combining the first signal component and the modified second signal component into an output audio signal, the output audio signal comprising two output channels configured for rendering by a headphone device.

[0015] In some, but not necessarily all, examples, the cross-channel mixing apparatus for cross-channel mixing a plurality of input channels comprises means for applying a head-related transfer function to each of the plurality of input channels before mixing the channels to produce a modified second signal component comprising two output channels, wherein the head-related transfer function applied to the input channel mixed to provide the output channel depends on an identity of the input channel and an identity of the output channel.

[0016] In some, but not necessarily all, examples, the cross-channel mixing apparatus for cross-channel mixing a plurality of input channels includes means for applying a headphone filter to each of the plurality of input channels before mixing the channels to produce a modified second signal component comprising two output channels, wherein the headphone filter applied to an input channel mixed to provide an output channel depends on an identity of the input channel and an identity of the output channel, wherein the headphone filter for an input channel mixes a direct version of the input channel with an ambient version of the input channel.

[0017] In some, but not necessarily all, examples, the relative gain of the direct version of the input channel compared to the ambient version of the input channel in the mix in the headphone filter is a user-controllable parameter.

[0018] In some, but not necessarily all, examples, the headphone filter for an input channel mixes a single-path direct version of the input channel with a multipath ambient version of the input channel; wherein a head-related transfer function is used to form the single-path direct version of the input channel; and wherein an indirect path filter is applied to each of the multipaths in combination with the head-related transfer function to form the multipath ambient version of the input channel. In some, but not necessarily all, examples, the indirect path filter includes a decorrelation device or a reverberation device.

[0019] In some, but not necessarily all, examples, the cross-channel mixing is configured to cause stereo widening of the headphone device such that a width of a spatial audio image associated with the modified second signal component is greater than a width of a spatial audio image associated with the second signal component prior to the cross-channel mixing of the second signal component.

[0020] In some but not necessarily all examples, the first portion is front and center relative to a user of the headphone device, and the second portion is peripheral relative to the user of the headphone device and does not overlap with the first portion.

[0021] In some, but not necessarily all, examples, the first portion and the second portion are contiguous.

[0022] In some, but not necessarily all, examples, the bypass means enables a component of the input audio signal representing a sound source that is coherent between the two stereo channels and located front and center to bypass the cross-channel mixing means.

[0023] In some, but not necessarily all, examples, the control input controls one or more of:

[0024] controlling the first portion and / or the second portion;

[0025] controlling the decomposition of an input signal into a first component and a second component;

[0026] controlling the relative gains of the first component and the second component;

[0027] controlling the widening of the second component;

[0028] controlling the direct to ambient gain ratio during the broadening of the second component;

[0029] Control the translation of the first component;

[0030] Controls whether there is translation of the first component;

[0031] controlling the translation range of the first component; and

[0032] Controls energy-based temporal smoothing.

[0033] In some but not necessarily all examples, when the input audio signal includes the same sound source repeated at different positions and rendered at the headphone device without interaural time difference and without frequency-related interaural level difference, when the sound source of the input audio signal is located at a first position that is front and center relative to a user of the headphone device, then when the sound source of the input audio signal is repeated at a second position, the sound source is rendered at the headphone device with interaural time difference and frequency-related interaural level difference, and the second position is relatively peripheral and not front and center of the user of the headphone device.

[0034] In some, but not necessarily all, examples, a system is provided that includes the device and a headphone device configured to receive and render the output audio signal.

[0035] In some, but not necessarily all, examples, the device is configured as a headphone device for rendering the output audio signal.

[0036] According to various, but not necessarily all, examples, there is provided a method for processing an input audio signal comprising at least one input channel / channels, the method comprising:

[0037] A first signal component including at least one input channel and a second signal component including a plurality of input channels are derived based on the input audio signal, wherein:

[0038] The first signal component depends on at least a first portion of a spatial audio image conveyed by the input audio signal, and the second signal component depends on at least a second portion of the spatial audio image different from the first portion;

[0039] cross-channel mixing at least some of the plurality of input channels of the second signal component to produce a modified second signal component while enabling the first signal component to bypass the cross-channel mixing; and

[0040] The first signal component and the modified second signal component are combined into an output audio signal comprising two output channels configured for rendering by a headphone device.

[0041] According to various, but not necessarily all, examples, there is provided an apparatus for processing an input audio signal comprising at least one input channel / channels, the apparatus comprising at least one processor; at least one memory comprising computer program code which, when executed by the at least one processor, causes the apparatus to:

[0042] A first signal component including at least one input channel and a second signal component including a plurality of input channels are derived based on the input audio signal, wherein:

[0043] The first signal component depends on at least a first portion of a spatial audio image conveyed by the input audio signal, and the second signal component depends on at least a second portion of the spatial audio image different from the first portion;

[0044] cross-channel mixing at least some of the plurality of input channels of the second signal component to produce a modified second signal component while enabling the first signal component to bypass the cross-channel mixing; and

[0045] The first signal component and the modified second signal component are combined into an output audio signal comprising two output channels configured for rendering by a headphone device.

[0046] According to various, but not necessarily all, examples, there is provided a computer program comprising computer readable program code configured to cause a computer to:

[0047] A first signal component comprising at least one input channel and a second signal component comprising a plurality of input channels are derived based on the input audio signal, wherein the first signal component depends on at least a first portion of a spatial audio image conveyed by the input audio signal and the second signal component depends on at least a second portion of the spatial audio image that is different from the first portion; and cross-channel mixing is performed on at least some of the plurality of input channels of the second signal component to produce a modified second signal component, while enabling the first signal component to bypass the cross-channel mixing.

[0048] According to various, but not necessarily all, examples, there is provided a device for processing an input audio signal comprising a plurality of channels to produce a two-channel output audio signal configured for rendering by a headphone device to produce a spatial audio image, the device comprising:

[0049] means for processing an input audio signal comprising a plurality of channels to produce a two-channel output audio signal configured for rendering by a headphone device;

[0050] Means for spatially processing the input audio signal to add position-dependent interaural time differences measurable between coherent audio events in two channels of the output audio signal and frequency-dependent and position-dependent interaural level differences measurable between coherent audio events in two channels of the output audio signal at peripheral positions rather than at central positions of the spatial audio image.

[0051] In some, but not necessarily all, examples, the means for deriving the first and second signal components is arranged to:

[0052] deriving the first signal component based on the input audio signal, the first signal component representing coherent sounds of the spatial audio image residing within the first portion of the spatial audio image; and

[0053] Based on the input audio signal, the second signal component is derived, the second signal component representing coherent sounds of the spatial audio image and incoherent sounds of the spatial audio image residing within the second portion of the spatial audio image and outside the first portion of the spatial audio image.

[0054] In some, but not necessarily all, examples, the first portion of the spatial audio image includes one or more angular ranges defining a set of sound arrival directions within the spatial audio image.

[0055] In some, but not necessarily all, examples, the one or more angular ranges include an angular range defining a range of sound arrival directions centered about a front direction of the spatial audio image.

[0056] In some but not necessarily all examples, the means for deriving the first and second signal components includes:

[0057] means for deriving respective coherence values for a plurality of frequency subbands based on the input audio signal, the coherence values describing a coherence between channels of the input audio signal in the respective frequency subbands;

[0058] means for deriving respective directional coefficients for the plurality of frequency subbands based on the estimated sound arrival directions according to the first portion of the spatial audio image, the respective directional coefficients indicating a relationship between the estimated sound arrival directions and the first portion of the spatial audio image in the respective frequency subbands;

[0059] means for deriving corresponding decomposition coefficients for the plurality of frequency subbands based on the coherence values and the directional coefficients; and

[0060] means for decomposing the input audio signal into the first and second signal components using the decomposition coefficients.

[0061] In some, but not necessarily all, examples, the means for deriving the directivity coefficient is arranged to, for the plurality of frequency subbands:

[0062] responsive to the estimated sound arrival direction of a frequency subband residing within the first portion of the spatial audio image, setting the directional coefficient for the frequency subband to a non-zero value; and

[0063] In response to the estimated sound arrival direction of a frequency subband residing within the second portion of the spatial audio image, the directional coefficient of the frequency subband is set to a value of zero.

[0064] In some but not necessarily all examples, the means for determining the decomposition coefficient is arranged to: for the plurality of frequency subbands, derive the respective decomposition coefficient as a product of the coherence value and a directional coefficient derived for the respective frequency subband.

[0065] In some but not necessarily all examples, the means for decomposing the input audio signal is arranged to, for the plurality of frequency sub-bands:

[0066] deriving a first signal component in each frequency subband as a product of the input audio signal in the corresponding frequency subband and a first scaling factor that increases with increasing values of the decomposition coefficients derived for the corresponding frequency subband; and

[0067] The second signal component in each frequency subband is derived as a product of the input audio signal in the corresponding frequency subband and a second scaling coefficient that decreases with increasing values of the decomposition coefficients derived for the corresponding frequency subband.

[0068] In some but not necessarily all examples, the apparatus includes means for delaying the first signal component by a predetermined time delay before combining the first signal component with the modified second signal component to create a delayed first signal component that is temporally aligned with the modified second signal component.

[0069] In some, but not necessarily all, examples, the apparatus includes means for modifying the first signal component before combining the first signal component with the modified second signal component, wherein the modification includes generating a modified first signal component based on the first signal component, wherein one or more sound source signals represented by the first signal component are shifted in the spatial audio image.

[0070] In some, but not necessarily all, examples, each of the plurality of input channels includes two channels.

[0071] According to various, but not necessarily exhaustive, examples are provided as claimed in the following claims.

[0072] According to an example embodiment, there is provided a computer program comprising a computer readable program code configured to cause the performance of at least the method according to the aforementioned example embodiments when said program code is executed on a computing device.

[0073] The computer program according to the exemplary embodiments may be embodied on a volatile or non-volatile computer-readable recording medium, for example, as a computer program product comprising at least one computer-readable non-transitory medium having program code stored thereon, which, when executed by a device, causes the device to perform at least the operations described above for the computer program according to the exemplary embodiments of the present invention.

[0074] The exemplary embodiments of the present invention presented in this patent application should not be interpreted as limiting the applicability of the appended claims. The verb "comprise" and its derivatives are used in this patent application as open limitations that do not exclude the presence of unrecited features. Unless expressly stated otherwise, the features described below can be freely combined with each other.

[0075] Certain features of the invention are set forth in the appended claims. However, aspects of the invention, both as to its construction and its method of operation, together with further objects and advantages thereof, will be best understood from the following description of certain exemplary embodiments when read in connection with the accompanying drawings.

[0076] definition

[0077] A headphone device is a device having a left audio output device worn at, above, or in the user's left ear and a right audio output device worn at, above, or in the user's right ear. The audio heard by the user in the left ear depends on the audio output by the left audio output device, and not on the audio output by the right audio output device. The audio heard by the user in the right ear depends on the audio output by the right audio output device, and not on the audio output by the left audio output device. The headphones receive input signals wirelessly or via a wired connection. In some, but not necessarily all, examples, the headphone device includes an acoustic isolator that isolates the user's ears from external ambient sounds. In some examples, the headphone device may include a "can" that covers the user's ear and provides at least some acoustic isolation. In some examples, the headphone device may include a deformable "bud" that fits tightly within the user's ear and provides at least some acoustic isolation. Each audio output device includes a transducer that converts a received electrical signal into sound pressure waves or vibrations.

[0078] Multi-channel audio signal: In this disclosure, the term multi-channel audio signal is used to refer to an audio signal having two or more channels.

[0079] Stereo signal: The term stereo signal is used to refer to a stereo audio signal.

[0080] Surround sound signal: The term surround sound is used to refer to a multi-channel audio signal with more than two channels. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Embodiments of the invention are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which:

[0082] Figure 1A A block diagram illustrating some elements of an audio processing system for headphones according to an example;

[0083] Figure 1B A block diagram illustrating some elements of an audio processing system for headphones according to an example;

[0084] Figure 2 A block diagram showing some elements of a device used to implement an audio processing system for headphones according to an example;

[0085] Figure 3 A block diagram illustrating some elements of a signal decomposer according to an example;

[0086] Figure 4 A block diagram illustrating some elements of a re-panner for headphones according to an example;

[0087] Figure 5 A block diagram illustrating some elements of a stereo widening processor for headphones according to an example;

[0088] Figure 6 A flowchart depicting a method of audio processing for headphones according to an example is shown; and

[0089] Figure 7 A block diagram illustrating some elements of a device according to an example is shown. DETAILED DESCRIPTION

[0090] In the following example, a device 100, 100', 50 for processing an input audio signal 101 comprising a plurality of channels is disclosed, the device 100, 100', 50 comprising: means 104 for deriving, based on the input audio signal 101, a first signal component 105-1 comprising at least one input channel and a second signal component 105-2 comprising a plurality of input channels, wherein the first signal component 105-1 depends on at least a first part of a spatial audio image conveyed by the input audio signal 101 and the second signal component 105-2 depends on at least a second part of the spatial audio image that is different from the first part; and cross-channel mixing means 112 for mixing the plurality of input channels across channels. , 112′; a device 104 for directing the second signal component 105-2 to a cross-channel mixing device 112, 112′ for cross-channel mixing at least some of the multiple input channels of the second signal component 105-2 to produce a modified second signal component 113, 113′; a bypass device 104, 106 for enabling the first signal component 105-1 to bypass the cross-channel mixing device 112, 112′; a device 114, 114′ for combining the first signal component 111, 111′ and the modified second signal component 113, 113′ into an output audio signal 115, which includes two output channels configured for rendering by the headphone device 20.

[0091] Figure 1A A block diagram of some components and / or entities of an audio processing system 100 is shown, which can serve as a framework for various embodiments of the audio processing techniques described in the present disclosure. The audio processing system 100 receives a stereo audio signal as an input signal 101 and provides a stereo audio signal with an at least partially widened spatial audio image as an output signal 115. The input signal 101 and the output signal 115 are referred to below as the stereo signal 101 and the widened stereo signal 115, respectively. In the following examples involving the audio processing system 100, each of these signals is assumed to be a corresponding two-channel stereo audio signal, unless explicitly stated otherwise. Furthermore, each intermediate audio signal derived based on the input signal 101 is also a corresponding two-channel audio signal, unless explicitly stated otherwise.

[0092] However, the audio processing system 100 can be readily generalized to a system capable of processing spatial audio signals (i.e., multi-channel audio signals having more than two channels, such as 5.1-channel spatial audio signals or 7.1-channel spatial audio signals), certain aspects of which will also be described in the examples provided below.

[0093] The audio processing system 100 may also receive control inputs 10 and indications 12 of target sound source (virtual loudspeaker) positions.

[0094] according to Figure 1A The illustrated example audio processing system 100 comprises: a transform entity (or transformer) 102 for transforming a stereo audio signal 101 from the time domain into a transform domain stereo signal 103; a signal decomposer 104 for deriving, based on the transform domain stereo signal 103, a first signal component 105-1 representing a focused portion of a spatial audio image and a second signal component 105-2 representing a non-focused portion of the spatial audio image; a re-panner 106 for generating, based on the first signal component 105-1, a modified first signal component 107, wherein one or more sound sources represented in the focused portion of the spatial audio image are repositioned according to a target configuration; and a re-panner 106 for transforming the modified first signal component 107 from the transform domain into the time domain. 09-1; an inverse transform entity 108-1 for transforming the second signal component 105-2 from the transform domain into a time domain second signal component 109-2; a delay unit 110 for delaying the modified first signal component 109-1 by a predetermined time delay; a stereo widening (for headphones) processor 112 for generating a modified second signal component 113 based on the second signal component 109-2, wherein the width of the spatial audio image is extended from the width of the second signal component 109-2; and a signal combiner 114 for combining the delayed first signal component 111 and the modified second signal component 113 into a widened stereo signal 115, which conveys a partially expanded spatial audio image.

[0095] Figure 1B A block diagram showing some components and / or entities of an audio processing system 100' is shown. Figure 1AA variation of the audio processing system 100 is shown. In the audio processing system 100′, the difference from the audio processing system 100 is that the inverse transform entities 108-1 and 108-2 are omitted, the delay element 110 is replaced by an optional delay element 110′ for delaying the modified first signal component 100 into a delayed modified first signal component 111′, the stereo widening processor 112 is replaced by a stereo widening processor 112′ for generating a modified (transform domain) second signal component 113′ based on the transform domain second signal component 105-2, and the signal combiner 114 is replaced by a signal combiner 114′ for combining the delayed modified first signal component 111′ and the modified second signal component 113′ into a widened stereo signal 115′ in the transform domain. In addition, the audio processing system 100′ comprises a transform entity 108′ for converting the widened stereo signal 115′ from the transform domain into a time-domain widened stereo signal 115. Where the optional delay element 110' is omitted, the signal combiner 114' receives the modified first signal component 107 (rather than a delayed version thereof) and operates to combine the modified first signal component 107 with the modified second signal component 113' to create a transform domain widened stereo signal 115'.

[0096] In the following, mainly through and according to Figure 1A The audio processing techniques described in the present disclosure are described with reference to the example audio processing system 100 and its entities, while the audio processing system 100′ and its entities are described separately when applicable. In further examples, the audio processing system 100 or the audio processing system 100′ may include further entities, and / or Figure 1A and 1B Some of the entities shown in may be omitted or combined with other entities. In particular, Figure 1A and Figure 1B and subsequently Figures 2 to 5 The logical components used to illustrate the corresponding entity, and therefore no structural limitations are applied to the implementation of the corresponding entity, but for example, the corresponding hardware device, the corresponding software device or the corresponding combination of the hardware device and the software device can be applied to implement any logical component of the entity separately from other logical components of the entity, to implement any sub-combination of two or more logical components of the entity, or to implement all logical components of the entity in combination.

[0097] The audio processing system 100, 100' can be implemented by one or more computing devices, and the resulting widened stereo signal 115 can be provided for playback via a headphone device. Generally, the audio processing system 100, 100' is implemented in any type of computing device, such as a portable handheld device, a desktop computer, a server device, etc. Examples of portable handheld devices include mobile phones, media player devices, tablet computers, laptop computers, etc. The computing device can also be used to play the widened stereo signal 115 through a headphone device. In another example, the audio processing system 100, 100' is provided in a headphone device, and playback of the widened stereo signal 115 is provided in the headphone device. In another example, a first portion of the audio processing system 100, 100' is provided in a first device, while a second portion of the audio processing system 100, 100' and playback of the widened stereo signal 115 are provided in the headphone device.

[0098] Figure 2 A block diagram of some components and / or entities of a portable handheld device 50 implementing an audio processing system 100 or an audio processing system 100′ is shown. For the sake of brevity and clarity of description, in the following description, it is assumed that the elements of the audio processing system 100, 100′ and the playback of the resulting widened stereo signal are provided in the device 50. The device 50 also includes a memory device 52 for storing information (e.g., the stereo signal 101), and a communication interface 54 for communicating with other devices and possibly receiving the stereo signal 101 therefrom. The device 50 optionally also includes an audio preprocessor 56 that can be used to preprocess the stereo signal 101 read from the memory 52 or received via the communication interface 54 before providing the stereo signal 101 to the audio processing system 100, 100′. For example, the audio preprocessor 56 can decode an audio signal stored in an encoded format into a time-domain stereo audio signal 101.

[0099] Still refer to Figure 2 The audio processing system 100 , 100 ′ may also receive a first control input 10 and an indication 12 together with the stereo signal 101 from or via the audio preprocessor 56 .

[0100] The control input 12 is used to control the signal decomposition 104 and / or the re-panning 106 and / or the stereo widening 112, 112'. More details are provided in the following description.

[0101] The position of the target sound source (virtual loudspeaker) is indicated by the indication 12. Effectively, this means the position of the loudspeaker if the input audio signal would be reproduced by the loudspeaker.

[0102] The virtual speaker positions are typically matched to the speaker format of the input audio signal. For a stereo input signal, the virtual speaker positions may, for example, correspond to speaker angles of + / - 30 degrees relative to the front direction. For a multi-channel audio signal, such as 5.1, these angles are typically 0, + / - 30, and + / - 110 degrees. However, in practice, the virtual speaker positions may have any meaningful values. The target sound source position indication may also be provided in other ways (via a user interface), may be a hard-coded value, or may be omitted. In at least some examples, the indication 12 is used to control signal decomposition 104. In some, but not necessarily all, examples, it may be used for stereo widening 112.

[0103] The audio processing system 100 , 100 ′ provides the widened stereo signal 115 derived therein to an interface for communicating with the headphone device 20 for rendering.

[0104] The headphone device 20 is a device having a left audio output device 21 worn at, above, or in a user's left ear, and a right audio output device 22 worn at, above, or in the user's right ear. The audio heard by the user in the left ear depends on the audio output by the left audio output device 21, not the audio output by the right audio output device 22. The audio heard by the user in the right ear depends on the audio output by the right audio output device 22, not the audio output by the left audio output device 21. The headphone device 20 receives input signals wirelessly or via a wired connection. In some, but not necessarily all, examples, the headphone device 20 includes acoustic isolators 23 that acoustically isolate the user's ears from the external environment. In some examples, the headphone device may include left and right "cans" 23 that cover the user's ears, house the respective audio output devices 21 and 22, and provide at least some acoustic isolation. In some examples, the headphone device may include deformable "buds" that fit snugly within the user's respective left and right ears, surrounding the respective audio output devices 21 and 22 and providing at least some acoustic isolation.

[0105] Each audio output device 21 , 22 includes a transducer that converts a received electrical signal into sound pressure waves or vibrations.

[0106] The stereo signal 101 can be received at the signal processing system 100, 100', for example, by reading the stereo signal from a memory or mass storage device in the device 50. In another example, the stereo signal is obtained via a communication interface (e.g., a network interface) from another device that stores the stereo signal in memory or from a mass storage device disposed therein. A widened stereo signal 115 can be provided for rendering by the headphone device 20. Additionally or alternatively, the widened stereo signal 115 can be stored in a memory or mass storage device in the device 50 and / or provided to another device via a communication interface for storage therein.

[0107] The information 12 defining the virtual speaker positions can be used to control the stereo widening process so that the audio source is perceived at a desired location, which may also be outside the physical location of the headphones. This process may include maintaining some portions (such as the focal portion of the spatial audio image) between the physical locations of the headphones.

[0108] The audio processing system 100, 100' can be arranged to process a stereo signal 101 arranged as a sequence of input frames, each input frame comprising a corresponding digital audio signal segment for each channel, provided as a corresponding time sequence of input samples at a predetermined sampling frequency. In a typical example, the audio processing system 100, 100' employs a fixed predetermined frame length. In other examples, the frame length can be a selectable frame length that can be selected from a plurality of predetermined frame lengths, or the frame length can be an adjustable frame length that can be selected from a predetermined range of frame lengths. The frame length can be defined as the number of samples L included in a frame for each channel of the stereo signal 101, which maps to a corresponding duration at the predetermined sampling frequency. By way of example, in this regard, the audio processing system 100, 100' can employ a fixed frame length of 20 milliseconds (ms), which yields frames of L = 160, L = 320, L = 640, and L = 960 samples per channel at sampling frequencies of 8, 16, 32, or 48 kHz, respectively. The frames can be non-overlapping or partially overlapping. However, these values are used as non-limiting examples, and frame lengths and / or sampling frequencies different from these examples may be used instead, depending on, for example, the required audio bandwidth, the required framing delay, and / or the available processing power.

[0109] Reference again Figure 1A and 1B, the audio processing system 100, 100' may include a transform entity 102 that is arranged to convert the stereo signal 101 from the time domain into a transform domain stereo signal 103. Typically, the transform domain relates to the frequency domain. In an example, the transform entity 102 uses a short-time discrete Fourier transform (STFT) to convert each channel of the stereo signal 101 into a corresponding channel of the transform domain stereo signal 103 using a predetermined analysis window length (e.g., 20 milliseconds). In another example, the transform entity 102 uses a (analysis) complex modulated quadrature mirror filter (QMF) library for time-frequency domain conversion. In this regard, the STFT and QMF libraries are used as non-limiting examples, and in further examples, any suitable transform technique known in the art may be used to create the transform domain stereo signal 103.

[0110] The transform entity 102 can further divide each channel into multiple frequency subbands, thereby obtaining a transform domain stereo signal 103, which provides a corresponding time-frequency representation for each channel of the stereo signal 101. The frequency bands in a given frame can be referred to as time-frequency tiles. The number of frequency subbands and the corresponding bandwidths of the frequency subbands can be selected, for example, based on the required frequency resolution and / or available computing power. In an example, the subband structure involves 24 frequency subbands based on the Bark scale, the equivalent rectangular band (ERB) scale, or the third octave scale known in the art. In other examples, a different number of frequency subbands with the same or different bandwidths can be used. A specific example in this regard is a single frequency subband that covers the entire input spectrum or a continuous subset thereof.

[0111] The time-frequency map block representing the frequency bin b in the time frame n of the channel i of the transform domain stereo signal 103 can be denoted as S(i, b, n). Channel i represents a single virtual loudspeaker or input channel. The transform domain stereo signal 103 (e.g., the time-frequency map block S(i, b, n)) is passed to the signal decomposer 104 to be decomposed into a first signal component 105-1 and a second signal component 105-2 therein. As mentioned above, a plurality of consecutive frequency bins can be grouped into a frequency subband, thereby providing a plurality of frequency subbands k=0,...,K-1. For each frequency subband k, the lowest bin (i.e., the frequency bin representing the lowest frequency in the frequency subband) can be denoted as b k,low , the highest bin (i.e., the frequency bin representing the highest frequency in the frequency sub-band) can be marked as b k,high .

[0112] Reference again Figure 1A and 1B, the audio processing system 100, 100' may include a signal decomposer 104, which is arranged to derive a first signal component 105-1 and a second signal component 105-2 based on the transform domain stereo signal 103. In the following, the first signal component 105-1 is referred to as a signal component representing a focused portion of the spatial audio image, and the second signal component 105-2 is referred to as a signal component representing a non-focused portion of the spatial audio image. The focused portion represents the portion of the audio image that is located in the front and center and can be regarded as the "front". The non-focused portion represents those portions of the audio image that are not represented by the focused portion (not the front and center) and can therefore be referred to as the "peripheral" portion of the spatial audio image. Here, the decomposition process does not change the number of channels, and therefore in this example, each of the first signal component 105-1 and the second signal component 105-2 is provided as a corresponding two-channel audio signal. It should be noted that the terms focused portion and non-focused portion used in this disclosure are names assigned to spatial sub-portions of the spatial audio image represented by the stereo signal 101, although these names do not imply any specific processing to be applied (or has been applied) to the base stereo signal 101 or the transform domain stereo signal 103, such as to actively emphasize or de-emphasize any portion of the spatial audio image represented by the stereo signal 101.

[0113] The signal decomposer 104 can derive a first signal component 105 based on the transform domain stereo signal 103, the first signal component 105 representing those coherent sounds of the spatial audio image that are within a predetermined focus range, and thus these sounds constitute the focused part of the spatial audio image. The focus range can be defined by the control input 10.

[0114] In contrast, the signal decomposer 104 can derive a second signal component 105 based on the transform domain stereo signal 103. The second signal component 105 represents coherent sound sources or sound components of the spatial audio image outside a predetermined focus range, as well as all incoherent sound sources of the spatial audio image. Such sound sources or components therefore constitute the non-focus portion of the spatial audio image. Thus, the signal decomposer 104 decomposes the sound field represented by the stereo signal 101 into a first signal component 105-1 that is excluded from the subsequent stereo widening process and a second signal component 105-2 that is subsequently subjected to the stereo widening process.

[0115] Figure 3 1 shows a block diagram of some components and / or entities of the signal decomposer 104 according to an example. Figure 3 As shown, the signal decomposer 104 can be conceptually divided into a decomposition analyzer 104a and a signal divider 126. Figure 3In other examples, the signal decomposer 104 may include further entities, and / or Figure 3 Some of the entities depicted may be omitted or combined with other entities.

[0116] The signal decomposer 104 may comprise a coherence analyzer 116 for estimating a coherence value 117 describing the coherence between channels of the transform domain stereo signal 103 based on the transform domain stereo signal 103. The coherence value 117 is provided to a decomposition coefficient determiner 124 for further processing therein.

[0117] The calculation of the coherence value 117 may involve deriving corresponding coherence values γ(k,n) for a plurality of frequency subbands k in a plurality of time frames n based on the time-frequency map S(i,b,n) representing the transform-domain stereo signal 103. As an example, the coherence value 117 may be calculated, for example, according to equation (3):

[0118]

[0119] Here, Re represents the real part operator and * represents the complex conjugate.

[0120] The term γ(k,n) is particularly valuable when the audio of a channel is dominated by an audio event that is common to both channels. The common audio event typically results in a complex phasor distribution across the entire frequency bin b. For all frequency bins within the band, in the case of perfect coherence (i.e., γ(k,n) = 1), the phases of the two channels are identical.

[0121] Still refer to Figure 3 The signal decomposer 104 may include an energy estimator 118 for estimating the energy of the transform domain stereo signal 103 based on the transform domain stereo signal 103. The energy value 119 is provided to the direction estimator 120 for use therein in direction angle estimation.

[0122] The calculation of the energy value 119 may involve deriving the corresponding energy values E(i, k, n) of the plurality of frequency subbands k in the plurality of audio channels i in the plurality of time frames n based on the time-frequency map tile S(i, b, n). As an example, the energy value E(i, k, n) may be calculated, for example, according to equation (4):

[0123]

[0124] Still refer to Figure 3The signal decomposer 104 may include a direction estimator 120 for estimating the perceived direction of arrival of the sound represented by the stereo signal 101 based on the energy values 119 according to the target virtual loudspeaker configuration applied in the stereo signal 101. The direction estimation may include calculating a direction angle 121 based on the energy values according to the target virtual loudspeaker positions, the direction angle 121 being provided to a focus estimator 122 for further analysis therein.

[0125] The target sound source (virtual speaker) configuration may also be referred to as the channel configuration (of the stereo signal 101). This information may be obtained, for example, from metadata 12 accompanying the stereo signal 101 (e.g. metadata included in the audio container in which the stereo signal 101 is stored). In another example, information defining the target virtual speaker configuration applied to the stereo signal 101 may be received (as user input) via a user interface of the device 50. The target virtual speaker configuration may be defined by indicating, for each channel of the stereo signal 101, a corresponding target virtual speaker position relative to an assumed listening point. As an example, the target position of the virtual speaker may include a target direction, which may be defined as an angle relative to a reference direction (e.g., a front direction). Thus, for example, in the case of a two-channel stereo signal, the target virtual speaker configuration may be defined as a corresponding target angle ∝ relative to the front direction of the left and right virtual speakers. in (1) and ∝ in (2). The target angle relative to the forward direction ∝ in (i) can alternatively be represented by a single target angle ∝ in Indicates that the single target angle defines the absolute value of the target angle relative to the forward direction, such that ∝ in (1)=∝ in And ∝ in (2)=-∝ in .

[0126] In a further example, no indication 12 is received in the audio processing system 100, 100' and predetermined information is applied instead in this regard by elements of the audio processing system 100, 100' (signal decomposer 104, re-panner 106) that define information for a target virtual speaker configuration to be applied in the stereo signal 101. An example in this regard involves applying a fixed predetermined target virtual speaker configuration. Another example involves selecting one of a plurality of predetermined target virtual speaker configurations based on the number of audio channels in the received stereo signal 101. Non-limiting examples in this regard include selecting a target virtual speaker configuration in which the channels are at ±30 degrees relative to the front direction in response to a two-channel signal 101 (which is therefore assumed to be a two-channel stereo audio signal), and / or selecting a target virtual speaker configuration in which the channels are at target angles ∝0, ±30 and ±110 degrees relative to the front direction in response to a six-channel signal (which is therefore assumed to represent a 5.1-channel surround sound signal). in (i) Target virtual loudspeaker configuration positioned and supplemented with a low-frequency effects (LFE) channel.

[0127] The direction estimator 120 is configured to estimate the perceived direction of arrival of the sound represented by the stereo signal 101. The direction estimation may involve estimating the direction of arrival of the sound represented by the stereo signal 101 based on the estimated energy E(i, k, n) and the target virtual speaker position ∝ in (i) deriving corresponding direction angles 121, θ(k,n), for a plurality of frequency subbands k in a plurality of time frames n, thereby indicating the estimated perceived direction of arrival of the sound in the frequency subband of the input frame. The direction estimation may be performed, for example, using the tangent law according to equations (5) and (6), where the underlying assumption is that the sound sources in the sound field represented by the stereo signal 101 are arranged in the desired spatial locations (to a significant extent) using amplitude translation:

[0128]

[0129] in

[0130]

[0131] Among them, ∝ in represents the target angles ∝ that define the target positions of the left and right virtual speakers relative to the front direction, respectively. in (1) and ∝ in (2), the left and right virtual speakers are positioned symmetrically (and equidistantly) relative to the front direction in this example. In other examples, the target positions of the left and right virtual speakers can be positioned asymmetrically relative to the front direction (e.g., such that |∝_in(1)|≠|∝_in(2)|). The modification of equation (5) solves this aspect, which is a simple task for those skilled in the art.

[0132] For example, in case of asymmetric (virtual) loudspeaker positions, a modification of equation (5) can be performed as follows. First, calculate half the angle between the loudspeakers:

[0133]

[0134] Next, calculate the midpoint between the speakers:

[0135]

[0136] Using these values, for the asymmetric case, equation (5) can be expressed as

[0137]

[0138] Where g1 and g2 are calculated in equation (6).

[0139] Still refer to Figure 3 , the signal decomposer 104 may include a focus estimator 122 for determining one or more focus coefficients 123 based on the estimated perceptible directions of arrival (direction angles 121) of the sounds represented by the stereo signal 101 according to a defined focus range within the spatial audio image, wherein the focus coefficients 123 indicate a relationship between the estimated directions of arrival (direction angles 121) of the sounds and the focus range. The focus range may, for example, be defined as a single angular range, or two or more angular sub-ranges, in the spatial audio image. In other words, the focus range may be defined as a set of directions of arrival of sounds within the spatial audio image. The focus range may be defined by the control input 10.

[0140] The focus estimator 122 may derive a focus coefficient 123 based at least in part on the direction angle 121. The focus estimator 122 may optionally also receive an indication 12 of a target virtual loudspeaker configuration to be used in the stereo signal 101 and further calculate the focus coefficient 123 based on this information. The focus coefficient 123 is provided to a decomposition coefficient determiner 124 for further processing therein.

[0141] Typically, one or more angular ranges of the focus range define a set of arrival directions that cover a defined portion around the center of the spatial audio image, rendering the focus estimate as a "front" estimate. Focus estimation can involve deriving corresponding focus (front) coefficients χ(k,n) for a plurality of frequency subbands k in a plurality of time frames n based on the direction angles 121θ(k,n), for example, according to equation (7):

[0142]

[0143] In equation (7), the first threshold θ Th1 and the second threshold θ Th2 , where θ Th1 <θ Th2 , which defines the main (central) angular focus range (angles around the front direction -θ Th1 to θ Th1 Between), auxiliary angle focus range (relative to the front direction from -θ Th2 to -θ Th1 and from θ Th1 to θ Th2 ) and non-focal range (relative to the front direction at -θ Th2 and θ Th2 outside). The coefficient θ that defines the focal range Th1 ,θ Th2 Can be defined by control input 10.

[0144] As a non-limiting example, the first threshold and the second threshold may be set to θ Th1 =5° and θ Th2 =15°, while in other examples, different thresholds θ may be used Th1 and θ Th2 Instead. Therefore, the focus estimation according to equation (7) applies a focus range including two angular ranges (i.e., a main angle focus range and a secondary angle focus range), and the focus coefficient χ(k, n) is set to unity in response to the sound source direction residing within the main angle focus range, and is set to zero in response to the sound source direction residing outside the focus range, while a predetermined function of the sound source direction is applied to set the focus coefficient χ(k, n) to a value between unity and 0 in response to the sound source direction residing within the secondary angle focus range. Typically, the focus coefficient χ(k, n) is set to a non-zero value in response to the sound source direction residing within the focus range, and is set to a zero value in response to the perceived sound source direction (direction angle 121θ(k, n)) residing outside the focus range. In an example, equation (7) may be modified such that the secondary angle focus range is not applied, and thus only a single threshold may be applied to define the limit between the focus range and the non-focus range.

[0145] Along the lines described above, the focus range may be defined as one or more consecutive non-overlapping angular focus ranges. As an example, the focus range may include a single defined angular range, or two or more defined angular ranges.

[0146] According to another example, at least one of the focus ranges is selectable, for example, such that an angular focus range can be selected or adjusted (e.g., by selecting or adjusting one or more thresholds defining the corresponding angular focus range) based on a target (or assumed) virtual loudspeaker configuration associated with the stereo input signal 12 and a focus range parameter present in the control input 10. For example, the control information can be used to control how much of the sound image (or what angle) is sent to widen.

[0147] Still refer to Figure 3 The signal decomposer 104 may include a decomposition coefficient determiner 124 for deriving decomposition coefficients 125 based on the coherence value 117 and the focus coefficient 123. The decomposition coefficients 125 are provided to a signal divider 126 for decomposing the transform domain stereo signal 103 therein.

[0148] The signal divider 126 is configured to derive a first signal component 105-1 representing a focused portion of the spatial audio image and a second signal component 105-2 representing a non-focused portion (e.g., a "peripheral" portion) of the spatial audio image based on the transform domain stereo signal 103 and the decomposition coefficients 125.

[0149] The decomposition coefficient determination is intended to provide a higher value for the decomposition coefficient β(k,n) for the frequency subband k and the frame n, which higher value indicates a higher coherence between the channels of the stereo signal 101 and conveys a directional sound component within the focus portion of the spatial audio image (see the description of the focus estimator 122 above). In this regard, the decomposition coefficient determination may involve deriving the corresponding decomposition coefficient β(k,n) for the plurality of frequency subbands k in the plurality of time frames n based on the corresponding coherence value γ(k,n) and the corresponding focus coefficient χ(k,n), for example, according to equation (8):

[0150] β(k,n)=γ(k,n)χ(k,n). (8)

[0151] In an example, the decomposition coefficients β(k,n) may be applied as decomposition coefficients 125 , eg provided to the signal divider 126 for decomposing the transform domain stereo signal 103 therein.

[0152] In another example, energy-based temporal smoothing is applied to the decomposition coefficients β(k,n) obtained from equation (8) in order to derive smoothed decomposition coefficients β′(k,n), which can be provided to the signal divider 126 to be applied thereto for decomposing the transform-domain stereo signal 103. The smoothing of the decomposition coefficients results in the sub-portions of the spatial audio image assigned to the first signal component 105-1 and the second signal component 105-2 changing more slowly over time, which can achieve improved perceptual quality in the resulting widened stereo signal by avoiding small-scale fluctuations in the spatial audio image therein. For example, a weighting providing energy-based temporal smoothing can be provided according to equation (9a):

[0153] β′(k,n)=A(k,n) / B(k,n), (9a)

[0154] in

[0155]

[0156] Wherein, E(k,n) represents the total energy of the transform domain stereo signal 103 of frequency subband k in time frame n (e.g., derivable based on the energy E(i,k,n) derived using equation (4)), and a and b (wherein, preferably, a+b=1) represent predetermined weighting factors. The weighting factors for the energy-based temporal smoothing (a and b) can be defined via the control input 10. As a non-limiting example, the values a=0.2 and b=0.8 can be applied, while in other examples, other values in the range of 0 to 1 can be applied instead.

[0157] Still refer to Figure 3 The signal decomposer 104 may include a signal divider 126 for deriving a first signal component 105-1 representing a focused portion of the spatial audio image and a second signal component 105-2 representing a non-focused portion (e.g., a “peripheral” portion) of the spatial audio image based on the transform domain stereo signal 103 and the decomposition coefficients 125.

[0158] As an example, the signal decomposition of multiple frequency subbands k in multiple channels i within multiple time frames n can be performed based on the time-frequency map S(i, b, n) according to equation (10a):

[0159]

[0160] Among them, S dr (i, b, n) denotes frequency bin b in time frame n of channel i of the first signal component 105-1 representing the focused portion of the spatial audio image,

[0161] S sw(i, b, n) denotes frequency bin b in channel i time frame n of the second signal component 105-2 of the non-focused portion (e.g. the “peripheral” portion) of the spatial audio image,

[0162] p represents a predetermined constant parameter (e.g., p=0.5 or 1), and

[0163] β(b,n) is equal to the decomposition coefficient β(k,n) for each frequency bin b within frequency subband k.

[0164] The signal divider 126 creates a first signal component 105-1 representing the focused portion of the spatial audio image and a second signal component 105-2 representing the non-focused portion (e.g., the "peripheral" portion) of the spatial audio image. However, it does not necessarily place the time-frequency map block S(i, b, n) into either the first signal component 105-1 or the second signal component 105-2. As in this example, it may scale or weight the contribution of the time-frequency map block S(i, b, n) more heavily in one of the first signal component 105-1 or the second signal component 105-2 depending on the decomposition coefficient β(k, n).

[0165] The scaling factor β(b,n) in equation (9) p can be replaced by another scaling factor that increases as the value of the decomposition coefficient β(b, n) increases (and decreases as the value of the decomposition coefficient β(b, n) decreases), and the scaling factor (1-β(b, n))p in equation (10a) can be replaced by another scaling factor that decreases as the value of the decomposition coefficient β(b, n) increases (and increases as the value of the decomposition coefficient β(b, n) decreases).

[0166] In another example, signal decomposition may be performed for multiple frequency subbands k in multiple channels i in multiple time frames n based on the time-frequency map S(i, b, n) according to equation (10b):

[0167]

[0168] Among them, β Th Represents a defined threshold value, whose value is in the range of 0 to 1, such as β Th =0.5. Signal decomposition parameter β Th can be defined by the control input 10. If equation (10b) is applied, the temporal smoothing of the decomposition coefficients 125 described above and / or the resulting signal component S sw (i, b, n) and S dr Temporal smoothing of (i, b, n) may be beneficial to improve the perceptual quality of the resulting widened stereo signal 115 .

[0169] The decomposition coefficients β(k, n) according to equation (8) are derived on a time-frequency tile basis, whereas equations (10a) and (10b) apply the decomposition coefficients β(b, n) on a frequency bin basis. In this regard, the decomposition coefficients β(k, n) derived for frequency subband k can be applied to each frequency bin b within frequency subband k.

[0170] Thus, the transform domain stereo signal 103 is divided in each time-frequency map tile S(i, b, n) into a first signal component 105-1 representing a sound component located in the focal portion of the spatial audio image represented by the stereo signal 101, and a second signal component 105-2 representing a sound component located outside the focal portion of the spatial audio image represented by the stereo signal 101. The first signal component 105-1 is then provided for playback without stereo widening applied thereto, while the second signal component 105-2 is then provided for playback after undergoing stereo widening.

[0171] Reference again Figure 1A and 1B The audio processing system 100, 100' may comprise a repanner 106 arranged to generate a modified first signal component 107 based on the first signal component 105-1, wherein one or more sound sources represented by the first signal component 105-1 are repositioned in the spatial audio image.

[0172] Figure 4 A block diagram of some components and / or entities of the retranslator 106 according to an example is shown. Figure 4 In other examples, the retranslator 106 may include further entities, and / or Figure 4 Some of the entities depicted in the figure may be omitted or combined with other entities.

[0173] The re-shifter 106 may include an energy estimator 128 for estimating the energy of the first signal component 105-1. The energy value 129 is provided to the direction estimator 130 and the re-shift gain determiner 136 for further processing therein. The energy value calculation may involve a time-frequency tile S dr (i, b, n) to derive the corresponding energy values E for multiple frequency subbands k in multiple audio channels i (multiple virtual speakers) in multiple time frames n dr (i, k, n). As an example, the energy value E can be calculated according to equation (11) dr (i, k, n):

[0174]

[0175] In another example, the energy value 119 calculated in the energy estimator 118 (e.g., according to equation (4)) can be reused in the re-panner 106, thereby eliminating the need for a dedicated energy estimator 128 in the re-panner 106. Even though the energy estimator 118 of the signal decomposer 104 estimates the energy value 119 based on the transform-domain stereo signal 103 instead of the first signal component 105-1, the energy value 119 enables the direction estimator 130 and the re-panning gain determiner 136 to operate correctly.

[0176] Still refer to Figure 4 The re-panner 106 may include a direction estimator 130 for estimating a perceived direction of arrival of a sound represented by the first signal component 105-1 based on the energy value 129 according to a target virtual loudspeaker configuration applied in the stereo signal 101. The direction estimation may include calculating a direction angle 131 based on the energy value 129 according to the target virtual loudspeaker position, the direction angle 131 being provided to a direction adjuster 132 for further processing therein.

[0177] The direction estimation may involve the energy E based on the estimation dr (i, k, n) and the position of the target virtual speaker ∝ in (i) derive the corresponding direction angles 131θ for multiple frequency subbands k in multiple time frames n dr (k, n). Direction angle 131θ dr (k, n) indicates the estimated perceived direction of arrival (direction angle 131) of the sound in the frequency subband of the first signal component 105-1. The direction estimation can be performed, for example, according to equations (12) and (13):

[0178]

[0179] in

[0180]

[0181] In another example, the direction angle 121 calculated in the energy estimator 128 (e.g., according to equations (5) and (6)) can be reused in the re-panner 106, thereby eliminating the need for a dedicated direction estimator 130 in the re-panner 106. Even if the direction estimator 120 of the signal decomposer 104 estimates the direction angle 121 based on the energy value 119 derived from the transform-domain stereo signal 103 instead of the first signal component 105-1, the sound source position angle is the same or substantially the same, and thus the direction angle 121 enables the direction adjuster 132 to operate correctly.

[0182] Still refer to Figure 4, the re-panner 106 may include a direction adjuster 132 for modifying the estimated perceived direction of arrival (direction angle 131) of the sound represented by the first signal component 105-1. The direction adjuster 132 may derive a modified direction angle 133 based on the direction angle 131. The modified direction angle 133 is provided to a panning gain determiner 134 for further processing therein.

[0183] Direction adjustment may include mapping a currently estimated perceived direction of arrival (direction angle 131 ) to a corresponding modified direction angle 133 according to the control information 10 , the corresponding modified direction angle 133 representing a new adjusted perceived direction of arrival of the sound.

[0184] A mapping between a currently estimated perceived direction of arrival (direction angle 131) and a new adjusted perceived direction of arrival (modified direction angle 132) may be provided by determining a mapping coefficient μ, which may be applied, for example, to derive corresponding modified direction angles θ′(k,n) for a plurality of frequency subbands k in a plurality of time frames n according to equation (15).

[0185] θ′(k,n)=μθ(k,n). (15)

[0186] The value of the mapping coefficient μ for the translation can be explicitly defined via the control input 10 .

[0187] If stereo widening 112 "widens" signal 105-2 by a certain amount, re-panner 106 widens signal 105-1 by re-panning by the same amount. As a practical example, stereo widening 112 may widen the signal so that a sound source that was originally at a position of 5 degrees is perceived after widening to be at a position corresponding to 10 degrees in the original signal. Thus, control information 10 may include information that re-panning by a factor of 2 (μ=2) is required so that the position of re-panned audio 107 matches the position of stereo widened audio 113.

[0188] Determining the mapping coefficient μ and deriving the modified direction angle θ′(k,n) according to equations (14) and (15) serves as a non-limiting example, and different procedures for deriving the modified direction angle 133 may alternatively be employed.

[0189] Still refer to Figure 4 , the re-panner 106 may include a panning gain determiner 134 for calculating a set of panning gains 135 based on the modified directional angle 133. The panning gain determination may include calculating respective panning gains g′(i,k,n) for a plurality of frequency subbands k in a plurality of audio channels i in a plurality of time frames n based on the modified directional angle θ′(k,n) using, for example, a Vector Basis Amplitude Panning (VBAP) technique known in the art.

[0190] For example, the translation gain g′(i, k, n) can be derived based on the tangent law:

[0191]

[0192]

[0193]

[0194]

[0195] Still refer to Figure 4 , the re-panner 106 may include a re-panning gain determiner 136 for deriving a re-panning gain 137 based on the panning gain 135 and the energy value 129. The re-panning gain 137 is provided to a re-panning processor 138 for deriving the modified first signal component 107 therein.

[0196] The re-translation gain determination process may include calculating the respective total energies E for a plurality of frequency sub-bands k in a plurality of time frames n, for example, according to equation (18). s (k, n):

[0197] E s (k, n) = ∑ i E dr (i, k, n). (18)

[0198] The re-translation gain determination may also include, for example, according to equation (19), based on the total energy E s (k, n) and the translation gain g′(i, k, n), calculate the corresponding target energy E for multiple frequency subbands k in multiple audio channels i in multiple time frames n t (i, k, n):

[0199] E t (i, k, n) = g′(i, k, n) 2 E s (k,n). (19)

[0200] The target energy E t (i, k, n) and energy value E dr (i, k, n) are applied together to derive corresponding re-panning gains for a plurality of frequency subbands k in a plurality of audio channels i in a plurality of time frames n, for example according to equation (20):

[0201]

[0202] In this example, the retranslation gain g is obtained from equation (20)r (i, k, n) may be applied as a re-translation gain 137, for example, which is provided to a re-translation processor 138 to derive therein the modified first signal component 107. In another example, energy-based temporal smoothing is applied to the re-translation gain g obtained from equation (20) r (i, k, n), to derive the smoothed re-translation gain g′ r (i, k, n), which may be provided to the re-translation processor 138 to be applied therein for re-translation. r Smoothing of (i, k, n) results in a slowing down of changes over time within the sub-portion of the spatial audio image assigned to the first signal component 105-1, which may improve the perceptible quality in the resulting widened stereo signal 115 by avoiding small-scale fluctuations in the corresponding portion of the widened spatial audio image.

[0203] Still refer to Figure 4 , the re-panner 106 may include a re-panning processor 138 for deriving a modified first signal component 107 based on the first signal component 105-1 in dependence on a re-panning gain 137. In the resulting modified first signal component 107, the sound source in the focus portion of the spatial audio image is repositioned (i.e., re-panned) according to the modified direction angle 132 derived in the direction adjuster 132 to account for (possible) differences between directly reproducing the stereo signal on headphones and reproducing the stereo signal processed by stereo widening 112 on headphones. The channels of the modified first signal component 107 are provided to an inverse transform entity 108-1 for conversion from the transform domain to the time domain therein.

[0204] The process of deriving the modified first signal component 107 may comprise, for example, depending on the re-translation gain g according to equation (21) r (i, b, n), based on the corresponding time-frequency block S of the first signal component 105-1 dr (i, b, n) derives the corresponding time-frequency map S for multiple frequency bins b in multiple audio channels i in multiple time frames n dr,rp (i, b, n):

[0205] S dr,rp (i, b, n) = g r (i, b, n)S dr (i, b, n). (21)

[0206] According to equation (20), the retranslation gain g r (i, k, n) is derived on a time-frequency patch basis, whereas Equation (21) applies the re-translation gain g on a frequency bin basis.r In this regard, the re-translation gain g derived for frequency sub-band k can be r (i, k, n) applies to each frequency bin b within frequency subband k.

[0207] In other examples, the translation may be applied to each time-frequency tile S(i, b, n) by applying a controlled gain g r Different combinations of (i, b, n), controlled reverberation or decorrelation and optionally controlled delays to produce the channels of the modified first signal component 107. Reverberation or decorrelation is typically added only at low sound levels.

[0208] In some embodiments, the modified first signal component 107 can be divided into two paths (e.g., using variables received in the control information 10). The signal in the second path is processed using reverberation or decorrelation. The signal in the first path is passed forward without processing and without any cross-channel mixing. The signals in the two paths are combined, for example, by summing them.

[0209] Return Reference Figure 1A , the audio processing system may comprise an inverse transform entity 108-1 arranged to transform (back) the channels of the modified first signal component 107 from the transform domain to the time domain, thereby providing a time domain modified first signal component 109-1. Along similar lines, the audio processing system 100 may comprise an inverse transform entity 108-2 arranged to transform (back) the channels of the second signal component 105-2 from the transform domain to the time domain, thereby providing a time domain second signal component 109-2. Both the inverse transform entity 108-1 and the inverse transform entity 108-2 utilize an applicable inverse transform that reverses the time to transform domain conversion performed in the transform entity 102. As a non-limiting example in this regard, the inverse transform entities 108-1, 108-2 may therefore apply an inverse STFT or a (synthetic) QMF library to provide the inverse transform. The resulting time domain modified first signal component 109-1 may be represented as s dr (i, m), and the resulting time-domain second signal component 109-2 can be expressed as s sw (i, m), where i represents a channel and m represents a time index (ie, a sample index).

[0210] Reference again Figure 1B As described above, in the audio processing system 100′, the inverse transform entities 108-1, 108-2 are omitted, and the modified first signal component 107 is provided as a transform domain signal to the (optional) delay element 110′, and the transform domain second signal component 105-2 is provided as a transform domain signal to the stereo widening processor 112′.

[0211] Return Reference Figure 1A , the audio processing system 100 may include a stereo widening processor 112 arranged to generate a modified second signal component 113 based on the second signal component 109-2, wherein the width of the spatial audio image is widened from the signal represented by the second signal component 109-2. The stereo widening processor 112 may apply any stereo widening technique known in the art to widen the width of the spatial audio image. In an example, the stereo widening processor 112 modifies the second signal component s sw (i, m) is processed into a modified second signal component s′ sw (i, m), where the second signal component s sw (i, m) and the modified second signal component s′ sw (i, m) are time domain signals respectively.

[0212] The stereo widening technique may involve adding a processed (e.g., filtered) version of a contralateral channel signal to each of the left and right channel signals of a stereo signal to derive an output stereo signal having a widened spatial audio image (the widened stereo signal). In other words, a processed version of the right channel signal of the stereo signal is added to the left channel signal of the stereo signal to create the left channel of the widened stereo signal, and a processed version of the left channel signal of the stereo signal is added to the right channel signal of the stereo signal to create the right channel of the widened stereo signal. The process of deriving the widened stereo signal may further involve pre-filtering (or otherwise processing) each of the left and right channel signals of the stereo signal before adding the corresponding processed contralateral signal to the stereo signal to preserve a desired frequency response in the widened stereo signal.

[0213] Along the above ideas, stereo widening is easily summarized as widening the spatial audio image of a multi-channel input audio signal, thereby deriving an output multi-channel audio signal (widened multi-channel signal) having a widened spatial audio image. In this regard, the processing involves creating the left channel of the widened multi-channel audio signal as the sum of the (first) filtered versions of the channels of the multi-channel input audio signal, and creating the right channel of the widened multi-channel audio signal as the sum of the (second) filtered versions of the channels of the multi-channel input audio signal. A dedicated predetermined filter can be provided for each pair of input channels (channels of the multi-channel input signal) and output channels (left and right). As an example in this regard, the left and right channel signals S of the widened multi-channel signal can be defined based on the channels of the multi-channel audio signal S, respectively, according to equation (1) out,left and S out,right :

[0214]

[0215] Where S(i, b, n) represents the frequency bin b in the time frame n of the channel i of the multi-channel signal S, H left (i, b) represents the frequency bin b for filtering channel i of the multi-channel signal S to create the left channel signal S out,left (b, n) of the corresponding channel components, and H right (i, b) represents the frequency bin b for filtering channel i of the multi-channel signal S to create the right channel signal S out,right (b, n) corresponding channel component filter. left (i, b) and H right (i, b) is a directional filter pair.

[0216] In stereo widening for headphones, filter H left (i, b) and H right (i, b) may include HRTFs, or HRTFs (or BRIR) may be used later in the processing chain. In stereo widening for headphones, the filter H left (i, b) can be a 90 degree HRTF (ie to the left). Filter H right (i, b) may be a HRTF of -90 degrees (ie to the right).

[0217] In stereo widening for headphones, filter H left (i, b) may comprise a direct (dry) portion and an ambient portion comprising one or more indirect (wet) paths.

[0218]

[0219] where r is the ratio between the direct portion and the ambient portion.

[0220] The direct ambient ratio r can be defined via the control input 10 .

[0221] Direct Partial Filter H left,direct (i, b) may be a HRTF of 90 degrees (ie, to the left).

[0222] For each time-frequency tile S(i, b, n), the indirect partial filter H left,ambient (i, b) can represent different indirect paths, each with controlled gain, controlled reverberation or decorrelation, and optionally controlled delay. Each different indirect path is processed using a corresponding HRTF. The directions of the HRTFs are usually chosen so that they cover several directions around the listener, creating an envelope and / or a sense of spaciousness. The filters of the different indirect paths are usually combined into a single filter H before application. left,ambient (i,b).

[0223] Likewise, the filter H right (i, b) may comprise a direct (dry) portion and an ambient portion comprising one or more indirect (wet) paths.

[0224]

[0225] where r is the ratio between the direct portion and the ambient portion.

[0226] Direct Partial Filter H right,direct (i, b) may be the HRTF for -90 degrees (ie, to the right).

[0227] For each time-frequency tile S(i, b, n), the indirect partial filter H right,ambient (i, b) can represent different indirect paths, each with controlled gain, controlled reverberation or decorrelation, and optionally controlled delay. Each different indirect path is processed using a corresponding HRTF. The directions of the HRTFs are usually chosen so that they cover several directions around the listener, creating an envelope and / or a sense of space. The filters of the different indirect paths are usually combined into a single filter H before application. right,ambient (i,b).

[0228] The target virtual loudspeaker position indication 12 may optionally be provided to a stereo widening block 112. The indicated virtual loudspeaker positions may then be used to provide a stereo widening block 112 for, for example, H left and H right The filter selects the corresponding HRTF, for example, for a stereo signal, a + / -30 degree HRTF is selected by default. However, to produce the strongest widening effect for a stereo signal, HRTFs up to + / -90 can be selected instead. In summary, the stereo widening block 112 can map the indicated virtual speaker positions to modified positions (to obtain a stronger widening effect), which are then used to derive the filter H left and H right .

[0229] Figure 5 Shown is a block diagram of some components and / or entities of the stereo widening processor 112 according to a non-limiting example.

[0230] The stereo widening processor 112 is configured to provide cross-channel mixing means for mixing the headphone filter H before mixing these channels to produce a modified second signal component 113 comprising two output channels (left and right). LL 、H RL 、H LR and H RR applied to each of a plurality of input channels, wherein a headphone filter H is applied to the input channels that are mixed to provide the output channels.mn Depends on the identity of the output channel m and the identity of the input channel n.

[0231] Headphone filter H mn A head-related transfer function may be included that depends on the identity of the output channel m and the identity of the input channel n.

[0232] Headphone filter H for input channel n mn The headphone filter may be configured to mix a directly rendered version of an input channel with an ambient rendered version of the input channel. In the mixing of the headphone filter, the relative gain of the direct version of the input channel compared to the ambient version of the input channel may be controlled by a user-controllable parameter r. The headphone filter of the input channel may be configured to mix a single-path direct version of the input channel with a multi-path ambient version of the input channel, wherein a head-related transfer function is used to form the single-path direct version of the input channel, and for each path in the multi-path, an indirect path filter and a head-related transfer function are used in combination to form the multi-path ambient version of the input channel. The indirect path filter may include a decorrelation device or a reverberation device.

[0233] The cross-channel mixing results in a stereo widening of the headphone device such that a width of a spatial audio image associated with the modified second signal component is larger than a width of a spatial audio image associated with the second signal component before the cross-channel mixing of the second signal component.

[0234] In this example, the four filters H LL 、H RL 、H LR and H RR is applied to create a widened spatial audio image: the left channel of the modified second signal component 113 is created as LL The left channel of the filtered second signal component 109-2 is mixed with the left channel of the second signal component 109-2 by the filter H LR The right channel of the filtered second signal component 109-2 is the sum of the right channels, while the right channel of the modified second signal component 113 is created as a result of the filter H RL The left channel of the filtered second signal component 109-2 is mixed with the left channel of the second signal component 109-2 by the filter H RR The right channel sum of the filtered second signal component 109-2. Figure 5 In the example of , the stereo widening process is performed based on the time domain second signal component 109-2. In other examples, the stereo widening process can be performed in the transform domain (e.g., using Figure 5 In this alternative example, the order of the inverse transform entity 108-2 and the stereo widening processor 112 is changed.

[0235] In an example, the stereo widening processor 112 may be provided with a dedicated filter bank H LL 、H RL 、H LR and H RR , which is designed to produce the desired degree of stereo widening for the target virtual speaker configuration. In another example, the stereo widening processor 112 may be provided with multiple sets of filters H LL 、H RL 、H LR and H RR , each set of filters is designed to produce a desired degree of stereo widening for a target virtual speaker configuration. In the latter example, a set of filters is selected according to the indicated target virtual speaker configuration. In the case of multiple sets of filters, the stereo widening processor 112 can dynamically switch between filter sets, for example, in response to changes in the indicated virtual speaker positions. There are several ways to design a set of filters H LL 、H RL 、H LR and H RR .

[0236] In stereo widening for headphones, filter H LL It can be the above filter H left (left, b), filter H LR It can be the above filter H left (riqht, b), filter H RR It can be the above filter H right (right, b), filter H RL It can be the above filter H right (left, b).

[0237] The stereo widening performed by the spatial audio processor 112 can be performed in the time domain ( Figure 1A ) or transform domain ( Figure 1B ) is executed in

[0238] Return Reference Figure 1A , the audio processing system 100 may include a delay element 110 arranged to delay the modified first signal component 109-1 by a predetermined time delay, thereby creating a delayed first signal component 111. The time delay is selected such that it matches or substantially matches the delay caused by the stereo widening process applied in the stereo widening processor 112, thereby keeping the delayed first signal component 111 aligned in time with the modified second signal component 113. In an example, the delay element 110 delays the modified first signal component s dr (i, m) is modified to the delayed first signal component s′ dr (i, m). Figure 1A In the example of , the time delay is applied in the time domain. In an alternative example, the order of the inverse transform entity 108-1 and the delay element 110 may be changed, resulting in a predetermined time delay being applied in the transform domain.

[0239] Reference again Figure 1B As previously mentioned, in audio processing system 100′, delay element 110′ is optional, and if included, is arranged to operate in the transform domain. In other words, a predetermined time delay is applied to modified first signal component 107 to create a delayed modified first signal component 111′ in the transform domain, which is provided to combiner signal 114′ as a transform domain signal. As will be appreciated from the foregoing, stereo widening 112 (using, for example, HRTFs) is required if the perception of sound sources outside the headphones is desired. However, sounds can be localized between the headphones without stereo widening; for example, re-panning can be used to localize sound sources between the headphones (sounds cannot be localized outside the headphones using this method). However, the focal portion only contains sounds near the center, so localizing them between the headphones is sufficient. Peripheral portion 113 may also contain sound sources perceived outside the headphone location. Focus portion 111 does not contain sound sources perceived outside the headphone location, but they can still be wider than their original size.

[0240] Return Reference Figure 1A , the audio processing system 100 may comprise a signal combiner 114 arranged to combine the delayed first signal component 111 and the modified second signal component 113 into a widened stereo signal 115, wherein the width of the spatial audio image is partially extended (in the periphery but not necessarily in the front focus part) from the width of the stereo signal 101. As an example in this regard, the widened stereo signal 115 may be derived as a sum, an average or another linear combination of the delayed first signal component 111 and the modified second signal component 113, e.g. according to equation (22):

[0241] s out (i, m) = s′ sw (i, m)+s′ dr (i, m), (22)

[0242] Among them, s out (i, m) represents the widened stereo signal 115 .

[0243] Reference again Figure 1BAs described above, in the audio processing system 100′, the signal combiner 114′ is arranged to operate in the transform domain, in other words to combine the (transform domain) delayed modified first signal component 113′ with the (transform domain) modified second signal component 113′ into a (transform domain) widened stereo signal 115′ to be provided to the inverse transform entity 108′. The inverse transform entity 108′ is arranged to convert the (transform domain) widened stereo signal 115′ from the transform domain into the (time domain) widened stereo signal 115. The transform entity 108′ may perform the conversion in a similar manner as described above in the context of the transform entities 108-1, 108-2.

[0244] Each of the exemplary audio processing systems 100, 100' described above by way of a number of examples can be further varied in a variety of ways. In the following, non-limiting examples in this regard are described.

[0245] In the above description of the elements of audio processing systems 100, 100', reference is made to processing of associated audio signals in a plurality of frequency subbands k. In one example, processing of the audio signals in each element of audio processing systems 100, 100' is performed across (all) frequency subbands k. In other examples, processing of the audio signals in at least some elements of audio processing systems 100, 100' is performed across a limited number of frequency subbands k. As examples, processing in a particular element of audio processing systems 100, 100' can be performed for a predetermined number of lowest frequency subbands k, for a predetermined number of highest frequency subbands k, or for a predetermined subset of frequency subbands k in the middle of a frequency range, such that a first predetermined number of lowest frequency subbands k and a second predetermined number of highest frequency subbands k are excluded from processing. Frequency subbands k excluded from processing (e.g., those at the lower end of the frequency range and / or those at the higher end of the frequency range) can be passed unmodified from the input to the output of the corresponding element. A non-limiting example of elements of the audio processing system 100, 100', in which processing may be performed only on a limited subset k of frequency subbands, involves one or both of the re-panner 116 and the stereo widening processor 112, 112', which may process the respective input signal only within a respective desired frequency subrange, for example in a predetermined number of lowest frequency subbands k or in a predetermined subset k of frequency subbands in the middle of the frequency range.

[0246] In another example, as already described above, the input audio signal 101 may include a multi-channel signal different from a two-channel stereo audio signal, such as a surround sound signal. For example, in the case where the input audio signal 101 includes a 5.1-channel surround sound signal, the audio processing techniques described above with reference to the left and right channels of the stereo signal 101 may be applied to the left and right front channels of the 5.1-channel surround sound signal to derive the left and right channels of the output audio signal 115. The other channels of the 5.1-channel surround sound signal may be processed, for example, so that they are multiplied by a predetermined gain factor (e.g., by a value having a value of The center channel of the 5.1 channel surround sound signal, scaled by a factor of 110 (e.g., 110 degrees relative to the front direction), is added to the left and right channels of the output audio signal 115 obtained from the audio processing system 100, 100′, while the left and right rear channels of the 5.1 channel surround sound signal may be processed using a conventional stereo widening technique using widening filters (using, for example, HRTF or BRIR) corresponding to respective target positions of the left and right rear speakers (e.g., ±110 degrees relative to the front direction). The LFE channel of the 5.1 channel surround sound signal may be added to the center signal of the 5.1 channel surround sound signal before the scaled version of the center signal of the 5.1 channel surround sound signal is added to the left and right channels of the output audio signal 115.

[0247] In another example, as previously described, the input audio signal 101 may include N spatially distributed channels that are processed to produce a two-channel audio signal 115 specifically for playback through a headphone device. Mixing the M channels to produce the first signal components 111, 111′ of the two-channel stereo audio signal 115 may occur at the re-panner 106. Mixing the M′ channels to produce the second signal components 113, 113′ of the two-channel stereo audio signal 115 may occur at a stereo widening processor of the headphone device 112.

[0248] An audio event (sound object) may move within the sound image. When the audio event (sound object) is within the focus range, the audio event is rendered using first signal components 111 and 111′ of a two-channel stereo audio signal 115. When the audio event is within the non-focus peripheral range, the audio event is rendered using second signal components 113 and 113′ of the two-channel stereo audio signal 115.

[0249] In another example, in addition or alternatively, the audio processing system 100, 100' can enable adjustment of the balance between the contributions from the first signal component 105-1 and the second signal component 105-2 in the resulting widened stereo signal 115. This can be provided, for example, by applying corresponding different scaling gains to the first signal component 105-1 (or a derivative thereof) and the second signal component 105-2 (or a derivative thereof). In this regard, the corresponding scaling gains can be applied, for example, in the signal combiner 114, 114' to scale the signal components derived from the first and second signal components 105-1, 105-2 accordingly, or to scale the first and second signal components 105-1, 105-2 accordingly in the signal divider 126. A single corresponding scaling gain can be defined for scaling the first and second signal components 105-1, 105-2 (or their respective derivatives) across all frequency subbands or in a predetermined subset of the frequency subbands. Alternatively or additionally, different scaling gains may be applied across frequency subbands, thereby enabling adjusting the balance between contributions from the first and second signal components 105-1, 105-2 only across certain frequency subbands and / or adjusting the balance differently across different frequency subbands.

[0250] In another example, alternatively or additionally, the audio processing system 100, 100' can enable scaling of one or both of the first signal component 105-1 and the second signal component 105-2 (or their respective derivatives) independently of one another, thereby enabling equalization (across frequency subbands) of one or both of the first and second signal components. This can be provided, for example, by applying corresponding equalization gains to the first signal component 105-1 (or its derivatives) and the second signal component 105-2 (or its derivatives). Dedicated equalization gains can be defined for one or more frequency subbands of the first signal component 105-1 and / or the second signal component 105-2. In this regard, for each of the first and second signal components 105-1, 105-2, an equalization gain can be applied, for example, in the signal divider 126 or in the signal combiner 114, 114', to scale the corresponding frequency subband of the corresponding one of the first and second signal components 105-1, 105-2 (or their respective derivatives). For a certain frequency sub-band, the equalization gains of both the first and second signal components 105-1, 105-2 may be the same, or different equalization gains may be applied to the first and second signal components 105-1, 105-2.

[0251] The operation of the audio processing systems 100 and 100' described above through various examples enables adaptive decomposition of a stereo signal 101 into a first signal component 105-1, representing a focused portion of a spatial audio image and provided for playback without applying stereo widening, and a second signal component 105-2, representing a peripheral (non-focused) portion of the spatial audio image that has been subjected to stereo widening. In particular, because the decomposition is performed based on the audio content conveyed frame by frame by frame in the stereo signal 101, the audio processing systems 100 and 100' can adapt to both relatively static spatial audio images having different characteristics and changes in the spatial audio image over time.

[0252] The disclosed stereo widening technology relies on excluding coherent sound sources within the focal portion of the spatial audio image from the stereo widening process and applying the stereo widening process primarily to coherent sounds and incoherent sounds (such as the environment) outside the focal portion, thereby improving sound quality and reducing the "coloration" of sounds within the focal portion while still providing a large degree of perceptible stereo widening.

[0253] In the previous example, the control input 10 can have one or more different functions:

[0254] The parameters of the decomposition process can be defined by control inputs. The control input 10 can, for example, define the focal range used in the analysis to divide the signal into focal (ie, fronto-central) and non-focal (ie, lateral) signals. The focal range can, for example, be defined by θ Th1 and θ Th2 or β Th To define the signal decomposition parameter β Th It can be defined, for example, via the control input 10 .

[0255] The control input 10 may, for example, control the relative gain between the widened peripheral signals 113, 113' and the non-widened front signals 111, 111'. For example, in some examples, it may control the relative gain ratio of the peripheral to the front.

[0256] The parameters of the widening process may be defined, for example, by a control input 10. The control input 10 may, for example, control the direct-to-ambient ratio r used in the widening. The parameters may include, for example, the direction in which the non-focused sounds are processed (e.g., with the aid of HRTF processing), and / or the amount of ambient (e.g., reverberation) or perceived externalization added to the sound to increase the "widening" effect. It is not necessary to process the non-focused sounds into different virtual directions, and an embodiment of the invention may enable the non-focused sounds to be processed using only reverberation, decorrelators, or other methods of increasing externalization of the non-focused sounds.

[0257] The control input 10 may, for example, explicitly or implicitly control whether panning occurs. For example, if the focus range is narrow, panning may not occur. For example, if the relative gain ratio of the periphery to the front is small, panning may not occur.

[0258] The value of the mapping coefficient μ that controls the degree of panning can, for example, be explicitly defined by a control input 10, or can be controlled by defining the focus range. An overpan factor μ can be used to modify the front center sector (i.e., the focus sound) in which the focus signal is perceived (e.g., it can be made to sound wider than the original signal). Control input 10 can also be another parameter or set of parameters that can modify where in the left and right panning dimensions the focus sound is heard.

[0259] The weighting factors for the energy-based temporal smoothing (a and b) may be defined, for example, by the control input 10 .

[0260] For example, all, some, or none of the control inputs may be controlled by user input.

[0261] The control input 10 may, for example, include parameters for controlling the focus sound (eg for adding ambience to produce a better externalization of the front sound).

[0262] The control input 10 may, for example, include parameters defining multiple analysis sectors (for decomposition) and multiple virtual speaker directions (for stereo widening blocks). Non-focused sounds may be divided into multiple sectors, not just left and right (outside the focus range). There may be several angular regions outside the focus range that can be processed separately, for example, to different directions or different amounts of ambience as described in the present invention.

[0263] The components of the audio processing system 100, 100' may be arranged, for example, according to Figure 6 The method 200 operates as a method for processing an input audio signal comprising a multi-channel audio signal representing a spatial audio image.

[0264] The method 200 includes:

[0265] At block 202: based on the input audio signal 101, a first signal component 105-1 comprising at least one input channel and a second signal component 105-2 comprising a plurality of input channels are derived, wherein

[0266] The first signal component 105 - 1 depends on at least a first (focused) part of the spatial audio image conveyed by the input audio signal 101 , and the second signal component 105 - 1 depends on at least a second (non-focused) part of the spatial audio image different from the first (focused) part.

[0267] The method 200 further includes, at block 204 , cross-channel mixing at least some of the plurality of input channels of the second signal component 105 - 2 to produce a modified second signal component 115 , while enabling the first signal component to bypass the cross-channel mixing.

[0268] The method 200 further includes, at block 206 , combining the first signal component 105 - 2 and the modified second signal component 115 into an output audio signal 115 comprising two output channels configured for rendering by a headphone device.

[0269] The method 200 may be varied in various ways, eg according to the examples described above in connection with the operation of the audio processing system 100 and / or the audio processing system 100 ′.

[0270] Cross-channel mixing enables the width of the spatial audio image to be extended from the width of the second signal component 105 - 2 .

[0271] Figure 7 A block diagram of some components of an exemplary device 300 is shown. The device 300 may include Figure 7 The device 300 may be employed in implementing one or more of the aforementioned components, for example, in the context of the audio processing system 100, 100'. The device 300 may implement, for example, the device 50 or one or more components thereof.

[0272] The device 300 comprises a processor 316 and a memory 315 for storing data and computer program code 317. The memory 315 and a portion of the computer program code 317 stored therein may further be arranged to, together with the processor 316, implement at least some of the operations, procedures and / or functions described above in the context of the audio processing systems 100, 100'.

[0273] The device 300 includes a communication section 312 for communicating with other devices. The communication section 312 includes at least one communication device that enables wired or wireless communication with other devices. The communication device of the communication section 312 may also be referred to as a corresponding communication means.

[0274] Device 300 may further include a user I / O (input / output) component 318, which may be arranged to provide a user interface, possibly in conjunction with processor 316 and a portion of computer program code 317, for receiving input from a user of device 300 and / or providing output to a user of device 300 to control at least some aspects of the operation of audio processing system 100, 100′ implemented by device 300. User I / O component 318 may include hardware components such as, for example, a display, a touch screen, a touchpad, a mouse, a keyboard, and / or an arrangement of one or more keys or buttons. User I / O component 318 may also be referred to as a peripheral device. Processor 316 may be arranged to control the operation of device 300, for example, based on a portion of computer program code 317 and possibly further based on user input received via user I / O component 318 and / or based on information received via communication portion 312.

[0275] Although processor 316 is depicted as a single component, it may be implemented as one or more separate processing components. Similarly, although memory 315 is depicted as a single component, it may be implemented as one or more separate components, some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cached storage.

[0276] The computer program code 317 stored in the memory 315 may include computer-executable instructions that, when loaded into the processor 316, control one or more operational aspects of the device 300. As an example, the computer-executable instructions may be provided as one or more sequences of one or more instructions. The processor 316 can load and execute the computer program code 317 by reading one or more sequences of one or more instructions included therein from the memory 315. The one or more sequences of one or more instructions may be configured to, when executed by the processor 316, cause the device 300 to perform at least some of the operations, processes, and / or functions described above in the context of the audio processing systems 100, 100′.

[0277] Thus, the device 300 may include at least one processor 316 and at least one memory 315, the memory 315 including computer program code 317 for one or more programs, the at least one memory 315 and computer program code 317 being configured, together with the at least one processor 316, to cause the device 300 to perform at least some of the operations, processes and / or functions described above in the context of the audio processing system 100, 100′.

[0278] The computer program stored in the memory 315 may be provided, for example, as a corresponding computer program product comprising at least one computer-readable non-transitory medium having computer program code 317 stored thereon, which, when executed by the device 300, causes the device 300 to perform at least some of the operations, processes, and / or functions described above in the context of the audio processing system 100, 100'. The computer-readable non-transitory medium may include a storage device or recording medium, such as a CD-ROM, DVD, Blu-ray disc, or another article of manufacture tangibly embodying the computer program. As another example, the computer program may be provided as a signal configured to reliably transmit the computer program.

[0279] References to processors should not be understood as covering only programmable processors but also special purpose circuits such as field programmable gate arrays (FPGAs), application specific circuits (ASICs), signal processors, etc. Features described above may be used in combinations other than those explicitly described.

[0280] In at least some of the foregoing examples, when the input audio signal 101 includes the same sound source repeated at different locations and rendered at the headphone device 20 without interaural time difference and without frequency-related interaural level difference, when the sound source of the input audio signal 101 is located at a first location that is relatively front and center of a user of the headphone device 30, then when the sound source of the input audio signal is repeated at a second location, the sound source is rendered at the headphone device 30 with interaural time difference and frequency-related interaural level difference, the second location being relatively peripheral and not front and center of the user of the headphone device 30.

[0281] The stereo widening processor 112, 112' (for headphones) spatially processes the input audio signal 101 to add position-dependent interaural time differences measurable between coherent audio events in the two channels of the output audio signal and frequency-dependent and position-dependent interaural level differences measurable between coherent audio events in the two channels of the output audio signal at peripheral positions rather than central positions of the spatial audio image.

[0282] In the aforementioned example, there is a bypass initiated by the signal decomposer 104 and provided via a bypass route including the re-panner 106, thereby enabling the first signal component 105-1 to bypass the stereo widening (for headphones) processor 112, 112'. In some, but not necessarily all, examples, the bypass enables components of the input audio signal 101 representing front and center sound sources that are coherent between the two stereo channels to bypass the cross-channel mix at the stereo widening processor 112, 112' (for headphones).

[0283] In at least some of the above examples, the first focal portion is front and center relative to the user of the headphone device, and the second portion is peripheral relative to the user of the headphone device. In at least some of the above examples, the first focal portion does not overlap with the first portion. In at least some of the above examples, the first focal portion and the second non-focal portion are continuous.

[0284] Although the above description discusses an embodiment in which there is a first focal portion and two second focal portions separated left and right by the first focal portion, other arrangements of the first and second focal portions are possible. Reference to a portion may, for example, refer to a single portion or a plurality of portions.

[0285] In case the second part comprises a plurality of parts, different spatial audio processing may be applied to each of the second parts. For example, different control inputs may be used for different second parts. The same control input may be used for different second parts that are arranged symmetrically on either side of a central direction. For example, different cross-channel mixes may be used for different second parts to achieve different widening effects. The same cross-channel mix may be used for different second parts that are arranged symmetrically on either side of a central direction. For example, different direct-to-ambient ratios r may be used for different second parts to achieve different effects. The same direct-to-ambient ratio r may be used for different second parts that are arranged symmetrically on either side of a central direction.

[0286] In case the first portion comprises a plurality of parts, then different processing, such as re-translation, may be applied to each of the second parts.

[0287] In the aforementioned example, the first (focus) portion remains fixed in the audio image when the headphone device is moved and the audio image is oriented relative to the headphone device. In other examples, the audio image is oriented relative to the "world" headphone device and is processed to rotate when the headphone device is rotated. In this example, the first (focus) portion can remain fixed in the audio image when the headphone device is moved, or can alternatively rotate with the headphone device. The headphone device 20 may include circuitry for tracking its orientation.

[0288] In some examples, the device 100, 100' is separate from the headphone device 20, e.g. Figure 3As shown. In other examples, the device 100, 100' is part of the headphone device 20. In at least some of the examples described above, the audio is divided into two paths, central and side sounds. For the central sound, the sound quality is important, so the processing is designed to keep this sound quality good. HRTF processing is avoided. The central sound can be widened, for example, by "re-panning", and although "re-panning" cannot produce a sound source outside the headphones, it does not degrade the sound quality and some widening is performed. For the side sounds, having the widest perception is most important. Therefore, HRTFs are used to achieve this effect (and the sound source is provided outside the headphones). This will degrade the sound quality, but this is a compromise to obtain the maximum width. Although one would maintain the sound quality for the central sound, it is better to widen them. Make the side sounds very wide.

[0289] Although some functions have been described above with reference to certain features and / or elements, those functions may be performed by other features and / or elements, whether described or not.Although features have been described with reference to certain embodiments, those features may also be present in other embodiments whether described or not.

Claims

1. An apparatus for processing at least one input audio signal of multi-channel audio, the apparatus comprising at least one processor and at least one memory, the at least one memory storing computer program code, the computer program code when executed by the at least one processor causing the apparatus to: Get control input for playing audio; Based on the obtained control input, a first signal component and a second signal component of the at least one input audio signal are derived, wherein the first signal component being dependent on at least a first portion of a spatial audio image conveyed by the at least one input audio signal, the second signal component being dependent on at least a second portion of the spatial audio image being different from the first portion, The first portion is a focused portion of the spatial audio image, the second portion is a non-focused portion of the spatial audio image and does not overlap with the first portion, and The first portion and the second portion are determined based at least in part on a coherence level between channels of the at least one input audio signal; localizing one or more coherent sound sources in the spatial audio image in the first portion; processing at least one of the first signal component and the second signal component based on the control input; as well as Based on the processed at least one of the first signal component and the second signal component, an output audio signal comprising at least two output channels is provided for rendering.

2. The device according to claim 1, wherein The processing of the first signal component is associated with one or more of the following: Gain, Pan, equalizing the first signal component without broadening or decorrelation of the first signal component, and The first signal component is externalized.

3. The device according to claim 1 or 2, wherein: The processing of the second signal component is associated with one or more of the following: Gain, HRTF processing, Decorrelation, HRTF processing and decorrelation, Additional equalization, and Pan.

4. The apparatus according to claim 1, wherein Processing of at least one of the first signal component and the second signal component is associated with one or more of: HRTF processing, Gain, position, balanced, widen, and Decorrelation.

5. The apparatus according to claim 1, wherein The control inputs control one or more of the following: controlling the first portion and / or the second portion; controlling the decomposition of an input signal into a first component and a second component; controlling the relative gains of the first component and the second component; controlling the widening of the second component; controlling the direct to ambient gain ratio during the broadening of the second component; Control the translation of the first component; Control whether there is translation of the first component; Control the translation range of the first component; as well as Controls energy-based temporal smoothing.

6. The apparatus according to claim 1, wherein The first signal component is a front center signal and the second signal component is a side signal.

7. The apparatus according to claim 1, wherein Processing at least one of the first and second signal components includes cross-channel mixing at least some of a plurality of input channels of the second signal component to produce a modified second signal component while enabling the first signal component to bypass the cross-channel mixing.

8. The apparatus according to claim 7, wherein Mixing the at least some of the plurality of input channels of the second signal component across channels comprises applying a head-related transfer function to each of the plurality of input channels before mixing the channels to produce a modified second signal component comprising two output channels, wherein the head-related transfer function applied to the input channels mixed to provide the output channels depends on identities of the input channels and identities of the output channels.

9. The apparatus according to claim 7 or 8, wherein Cross-channel mixing of at least some of the plurality of input channels of the second signal component comprises applying a headphone filter to each of the plurality of input channels before mixing the channels to produce a modified second signal component comprising two output channels, wherein the headphone filter applied to the input channel mixed to provide the output channel depends on an identity of the input channel and an identity of the output channel, wherein the headphone filter for the input channel mixes a direct version of the input channel with an ambient version of the input channel.

10. The apparatus according to claim 9, wherein In the mix in the headphone filter, the relative gain of the direct version of the input channel compared to the ambient version of the input channel is a user controllable parameter.

11. The apparatus according to claim 9, wherein The headphone filter for an input channel mixes a single-path direct version of the input channel with a multi-path ambient version of the input channel; and wherein a head-related transfer function is used to form the single-path direct version of the input channel; wherein an indirect path filter is combined with a head-related transfer function for each of the multipaths to form the multipath ambient version of the input channel.

12. The apparatus according to claim 11, wherein The indirect path filter comprises a decorrelation device or a reverberation device.

13. The apparatus according to claim 7, wherein The cross-channel mixing results in a stereo widening of the headphone device such that a width of a spatial audio image associated with the modified second signal component is larger than a width of a spatial audio image associated with the second signal component before the cross-channel mixing of the second signal component.

14. The apparatus according to claim 7, wherein The first portion and the second portion are continuous.

15. The apparatus according to claim 1, wherein When the at least one input audio signal includes the same sound source repeated at different positions and rendered without interaural time difference and without frequency-related interaural level difference at the headphone device, when the sound source of the at least one input audio signal is located at a first position that is relatively front and center of a user of the headphone device, then when the sound source of the at least one input audio signal is repeated at a second position, the sound source is rendered with interaural time difference and frequency-related interaural level difference at the headphone device, and the second position is relatively peripheral and not front and center of the user of the headphone device.

16. The device of claim 1, configured as a headphone device for rendering the output audio signal.

17. The apparatus according to claim 16, wherein The headphone device is configured to generate a spatial audio image, and the computer program code, when executed by the at least one processor, further causes the device to: processing the at least one input audio signal to produce a two-channel output audio signal configured for rendering; The at least one input audio signal is spatially processed to add position-dependent interaural time differences measurable between coherent audio events in two channels of the output audio signal and frequency-dependent and position-dependent interaural level differences measurable between coherent audio events in two channels of the output audio signal at peripheral positions rather than at central positions of the spatial audio image.

18. A method for processing at least one input audio signal of multi-channel audio, the method comprising: Get control input for playing audio; Based on the obtained control input, a first signal component and a second signal component of the at least one input audio signal are derived, wherein the first signal component being dependent on at least a first portion of a spatial audio image conveyed by the at least one input audio signal, the second signal component being dependent on at least a second portion of the spatial audio image being different from the first portion, The first portion is a focused portion of the spatial audio image, the second portion is a non-focused portion of the spatial audio image and does not overlap with the first portion, and The first portion and the second portion are determined based at least in part on a coherence level between channels of the at least one input audio signal; localizing one or more coherent sound sources in the spatial audio image in the first portion; processing at least one of the first signal component and the second signal component based on the control input; and Based on the processed at least one of the first signal component and the second signal component, an output audio signal comprising at least two output channels is provided for rendering.